Correcting biased value estimation in mixing value-based multi-agent reinforcement learning by multiple choice learning. (November 2022)
- Record Type:
- Journal Article
- Title:
- Correcting biased value estimation in mixing value-based multi-agent reinforcement learning by multiple choice learning. (November 2022)
- Main Title:
- Correcting biased value estimation in mixing value-based multi-agent reinforcement learning by multiple choice learning
- Authors:
- Liu, Bing
Xie, Yuxuan
Feng, Lei
Fu, Ping - Abstract:
- Abstract: Multi-agent reinforcement learning (MARL) has become more and more popular over the past decades, and many value-based MARL methods are proposed in the past few years. Neural networks play important roles in these methods and are used to predict the value of the state–action pair, i.e. Q-value and actions of agents are chosen based on this. However the inaccurate prediction of the neural network leads to the biased Q-value estimation, which will cause inefficient usage of the experience data and poor performance. Unlike ensemble methods that just reduce the variance of predictions, multiple choice learning (MCL) methods exploit the cooperation among all the candidate models. This paper corrects the biased Q-value by exploiting the collaboration between the ensemble model and MARL to obtain a stabler and preciser Q-value estimator. In this paper, a new MARL method called Multiple Choice QMIX is developed to address the biased Q-value issue, which also extends the application scenarios of MCL methods. Specifically, we propose a voting network to learn the confidence level of each estimator and thus can provide the best prediction by combining their results. And a voting hindsight loss is proposed to encourage the voting network to overcome the overestimation of the Q-value. We also conduct experiments on four challenging tasks of the StarCraft II micromanagement benchmark. Experiment results show that our method obtains a faster convergence rate and stablerAbstract: Multi-agent reinforcement learning (MARL) has become more and more popular over the past decades, and many value-based MARL methods are proposed in the past few years. Neural networks play important roles in these methods and are used to predict the value of the state–action pair, i.e. Q-value and actions of agents are chosen based on this. However the inaccurate prediction of the neural network leads to the biased Q-value estimation, which will cause inefficient usage of the experience data and poor performance. Unlike ensemble methods that just reduce the variance of predictions, multiple choice learning (MCL) methods exploit the cooperation among all the candidate models. This paper corrects the biased Q-value by exploiting the collaboration between the ensemble model and MARL to obtain a stabler and preciser Q-value estimator. In this paper, a new MARL method called Multiple Choice QMIX is developed to address the biased Q-value issue, which also extends the application scenarios of MCL methods. Specifically, we propose a voting network to learn the confidence level of each estimator and thus can provide the best prediction by combining their results. And a voting hindsight loss is proposed to encourage the voting network to overcome the overestimation of the Q-value. We also conduct experiments on four challenging tasks of the StarCraft II micromanagement benchmark. Experiment results show that our method obtains a faster convergence rate and stabler performance in multi-agent tasks. … (more)
- Is Part Of:
- Engineering applications of artificial intelligence. Volume 116(2022)
- Journal:
- Engineering applications of artificial intelligence
- Issue:
- Volume 116(2022)
- Issue Display:
- Volume 116, Issue 2022 (2022)
- Year:
- 2022
- Volume:
- 116
- Issue:
- 2022
- Issue Sort Value:
- 2022-0116-2022-0000
- Page Start:
- Page End:
- Publication Date:
- 2022-11
- Subjects:
- Dec-POMDPs -- Multi-agent reinforcement learning -- Multiple choice learning -- Biased value estimation
Engineering -- Data processing -- Periodicals
Artificial intelligence -- Periodicals
Expert systems (Computer science) -- Periodicals
Ingénierie -- Informatique -- Périodiques
Intelligence artificielle -- Périodiques
Systèmes experts (Informatique) -- Périodiques
Artificial intelligence
Engineering -- Data processing
Expert systems (Computer science)
Periodicals
620.00285 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09521976 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.engappai.2022.105329 ↗
- Languages:
- English
- ISSNs:
- 0952-1976
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3755.704500
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 24158.xml