Dynamic regret convergence analysis and an adaptive regularization algorithm for on-policy robot imitation learning. (September 2021)
- Record Type:
- Journal Article
- Title:
- Dynamic regret convergence analysis and an adaptive regularization algorithm for on-policy robot imitation learning. (September 2021)
- Main Title:
- Dynamic regret convergence analysis and an adaptive regularization algorithm for on-policy robot imitation learning
- Authors:
- Lee, Jonathan N.
Laskey, Michael
Tanwani, Ajay Kumar
Aswani, Anil
Goldberg, Ken - Other Names:
- Morales Marco guest-editor.
Tapia Lydia guest-editor.
Sánchez-Ante Gildardo guest-editor.
Hutchinson Seth guest-editor. - Abstract:
- On-policy imitation learning algorithms such as DAgger evolve a robot control policy by executing it, measuring performance (loss), obtaining corrective feedback from a supervisor, and generating the next policy. As the loss between iterations can vary unpredictably, a fundamental question is under what conditions this process will eventually achieve a converged policy. If one assumes the underlying trajectory distribution is static (stationary), it is possible to prove convergence for DAgger. However, in more realistic models for robotics, the underlying trajectory distribution is dynamic because it is a function of the policy. Recent results show it is possible to prove convergence of DAgger when a regularity condition on the rate of change of the trajectory distributions is satisfied. In this article, we reframe this result using dynamic regret theory from the field of online optimization and show that dynamic regret can be applied to any on-policy algorithm to analyze its convergence and optimality. These results inspire a new algorithm, Adaptive On-Policy Regularization (Aor ), that ensures the conditions for convergence. We present simulation results with cart–pole balancing and locomotion benchmarks that suggestAor can significantly decrease dynamic regret and chattering as the robot learns. To the best of the authors' knowledge, this is the first application of dynamic regret theory to imitation learning.
- Is Part Of:
- International journal of robotics research. Volume 40:Number 10/11(2021)
- Journal:
- International journal of robotics research
- Issue:
- Volume 40:Number 10/11(2021)
- Issue Display:
- Volume 40, Issue 10/11 (2021)
- Year:
- 2021
- Volume:
- 40
- Issue:
- 10/11
- Issue Sort Value:
- 2021-0040-NaN-0000
- Page Start:
- 1284
- Page End:
- 1305
- Publication Date:
- 2021-09
- Subjects:
- Imitation learning -- online optimization -- online learning -- dynamic regret
Robots -- Periodicals
Robots, Industrial -- Periodicals
629.89205 - Journal URLs:
- http://ijr.sagepub.com/ ↗
http://www.uk.sagepub.com/home.nav ↗ - DOI:
- 10.1177/0278364920985879 ↗
- Languages:
- English
- ISSNs:
- 0278-3649
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 16982.xml