Beyond backpropagate through time: Efficient model‐based training through time‐splitting. Issue 10 (24th May 2022)
- Record Type:
- Journal Article
- Title:
- Beyond backpropagate through time: Efficient model‐based training through time‐splitting. Issue 10 (24th May 2022)
- Main Title:
- Beyond backpropagate through time: Efficient model‐based training through time‐splitting
- Authors:
- Gao, Jiaxin
Guan, Yang
Li, Wenyu
Li, Shengbo Eben
Ma, Fei
Zheng, Jianfeng
Wei, Junqing
Zhang, Bo
Li, Keqiang - Abstract:
- Abstract: Model‐based policy gradient (MBPG) has been employed to seek an approximate solution to the optimal control problem. However, there is coupling between adjacent states due to temporal dependencies, making the training time grow linearly with the time horizon. This paper reshapes the training process of MBPG with the time‐splitting technique to establish a time‐independent algorithm called Training Through Time‐Splitting (T3S). First, copy the coupled variables to obtain two independent variables. Meanwhile, an extra variable together with an equivalence constraint is introduced for problem consistency. Then, the transformed problem divides into subproblems with carefully derived loss functions. Subproblems own decoupled variables and shared policy networks, which means they can be optimized concurrently. Guided by the algorithm design, this paper further proposes an asynchronous parallel training scheme to accelerate training efficiency. Numerical simulation shows that the T3S algorithm outperforms the MBPG algorithm by 83.6% in wall‐clock time with a trajectory tracking task.
- Is Part Of:
- International journal of intelligent systems. Volume 37:Issue 10(2022)
- Journal:
- International journal of intelligent systems
- Issue:
- Volume 37:Issue 10(2022)
- Issue Display:
- Volume 37, Issue 10 (2022)
- Year:
- 2022
- Volume:
- 37
- Issue:
- 10
- Issue Sort Value:
- 2022-0037-0010-0000
- Page Start:
- 8046
- Page End:
- 8067
- Publication Date:
- 2022-05-24
- Subjects:
- model‐based policy gradient -- optimal control -- parallel training -- reinforcement learning -- time‐splitting
Artificial intelligence -- Periodicals
Expert systems (Computer science) -- Periodicals
Intelligence artificielle -- Périodiques
Systèmes experts (Informatique) -- Périodiques
006.3 - Journal URLs:
- http://onlinelibrary.wiley.com/journal/10.1002/(ISSN)1098-111X ↗
https://www.hindawi.com/journals/ijis ↗
http://onlinelibrary.wiley.com/ ↗ - DOI:
- 10.1002/int.22928 ↗
- Languages:
- English
- ISSNs:
- 0884-8173
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 4542.310500
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 23231.xml