Off-Policy Q-Learning for Anti-Interference Control of Multi-Player Systems⁎Research supported by the National Natural Science Foundation of China under Grants 61673280, 61673100, the Open Project of Key Field Alliance of Liaoning Province under Grant 2019-KF-03-06 and the Project of Liaoning Shihua University under Grant 2018XJJ-005. Issue 2 (2020)
- Record Type:
- Journal Article
- Title:
- Off-Policy Q-Learning for Anti-Interference Control of Multi-Player Systems⁎Research supported by the National Natural Science Foundation of China under Grants 61673280, 61673100, the Open Project of Key Field Alliance of Liaoning Province under Grant 2019-KF-03-06 and the Project of Liaoning Shihua University under Grant 2018XJJ-005. Issue 2 (2020)
- Main Title:
- Off-Policy Q-Learning for Anti-Interference Control of Multi-Player Systems⁎Research supported by the National Natural Science Foundation of China under Grants 61673280, 61673100, the Open Project of Key Field Alliance of Liaoning Province under Grant 2019-KF-03-06 and the Project of Liaoning Shihua University under Grant 2018XJJ-005.
- Authors:
- Li, Jinna
Xiao, Zhenfei
Chai, Tianyou
Lewis, Frank. L.
Jagannathan, Sarangapani - Abstract:
- Abstract: This paper develops a novel off-policy game Q-learning algorithm to solve the anti-interference control problem for discrete-time linear multi-player systems using only data without requiring system matrices to be known. The primary contribution of this paper lies in that the Q-learning strategy employed in the proposed algorithm is implemented in an off-policy policy iteration approach other than on-policy learning due to the well-known advantages of off-policy Q-learning over on-policy Q-learning. All of the players work hard together for the goal of minimizing their common performance index meanwhile defeating the disturbance that tries to maximize the specific performance index, and finally they reach the Nash equilibrium of the game resulting in satisfying disturbance attenuation condition. In order to find the solution to the Nash equilibrium, the anti-interference control problem is first transformed into an optimal control problem. Then an off-policy Q-learning algorithm is proposed in the framework of typical adaptive dynamic programming (ADP) and game architecture, such that control policies of all players can be learned using only measured data. Comparative simulation results are provided to verify the effectiveness of the proposed method.
- Is Part Of:
- IFAC-PapersOnLine. Volume 53:Issue 2(2020)
- Journal:
- IFAC-PapersOnLine
- Issue:
- Volume 53:Issue 2(2020)
- Issue Display:
- Volume 53, Issue 2 (2020)
- Year:
- 2020
- Volume:
- 53
- Issue:
- 2
- Issue Sort Value:
- 2020-0053-0002-0000
- Page Start:
- 9189
- Page End:
- 9194
- Publication Date:
- 2020
- Subjects:
- H∞ control -- off-policy Q-learning -- game theory -- Nash equilibrium
Automatic control -- Periodicals
629.805 - Journal URLs:
- https://www.journals.elsevier.com/ifac-papersonline/ ↗
http://www.sciencedirect.com/ ↗ - DOI:
- 10.1016/j.ifacol.2020.12.2180 ↗
- Languages:
- English
- ISSNs:
- 2405-8963
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 23658.xml