Dynamic regret convergence analysis and an adaptive regularization algorithm for on-policy robot imitation learning. (September 2021)