Whittle index based Q-learning for restless bandits with average reward. (May 2022)
- Record Type:
- Journal Article
- Title:
- Whittle index based Q-learning for restless bandits with average reward. (May 2022)
- Main Title:
- Whittle index based Q-learning for restless bandits with average reward
- Authors:
- Avrachenkov, Konstantin E.
Borkar, Vivek S. - Abstract:
- Abstract: A novel reinforcement learning algorithm is introduced for multiarmed restless bandits with average reward, using the paradigms of Q-learning and Whittle index. Specifically, we leverage the structure of the Whittle index policy to reduce the search space of Q-learning, resulting in major computational gains. Rigorous convergence analysis is provided, supported by numerical experiments. The numerical experiments show excellent empirical performance of the proposed scheme.
- Is Part Of:
- Automatica. Volume 139(2022)
- Journal:
- Automatica
- Issue:
- Volume 139(2022)
- Issue Display:
- Volume 139, Issue 2022 (2022)
- Year:
- 2022
- Volume:
- 139
- Issue:
- 2022
- Issue Sort Value:
- 2022-0139-2022-0000
- Page Start:
- Page End:
- Publication Date:
- 2022-05
- Subjects:
- Discrete event system -- Reinforcement learning -- Restless bandits -- Whittle index -- Q-learning -- Average reward
Automatic control -- Periodicals
Automation -- Periodicals
629.805 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00051098 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.automatica.2022.110186 ↗
- Languages:
- English
- ISSNs:
- 0005-1098
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 1829.450000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 21101.xml