Transferring knowledge from human-demonstration trajectories to reinforcement learning. (January 2018)
- Record Type:
- Journal Article
- Title:
- Transferring knowledge from human-demonstration trajectories to reinforcement learning. (January 2018)
- Main Title:
- Transferring knowledge from human-demonstration trajectories to reinforcement learning
- Authors:
- Wang, Guo-fang
Fang, Zhou
Li, Ping
Li, Bo - Abstract:
- Nowadays, transfer learning (TL) has become a crucial technique to accelerate the slow optimization procedure of reinforcement learning (RL) by re-utilizing knowledge acquired in a previous related task. Nevertheless, most of the current relevant research acquires knowledge through RL training in the source task, which would be too time-consuming. In view of this situation, in this paper, we propose a novel TL framework where the agent extracts knowledge from human-demonstration trajectories of the source task and reuses the knowledge in RL in the target task. As for what to transfer, two forms of knowledge deduced from the demonstration trajectories, which are the k -nearest neighbour of the current state in source samples and visit frequency of homologous states, are adopted. For how to transfer, the two forms of knowledge are respectively used to recommend a preferred action when random exploration is needed and to shape an instantaneous reward for RL. Simulation experiments of balancing Cart-Poles with different difficulties suggest that both the two forms of knowledge accelerate the learning process of RL obviously. What is more, the effect is even more significant when they are used in combination. In this case, the experimental results manifest the positive role of our framework in RL.
- Is Part Of:
- Transactions of the Institute of Measurement and Control. Volume 40:Number 1(2018)
- Journal:
- Transactions of the Institute of Measurement and Control
- Issue:
- Volume 40:Number 1(2018)
- Issue Display:
- Volume 40, Issue 1 (2018)
- Year:
- 2018
- Volume:
- 40
- Issue:
- 1
- Issue Sort Value:
- 2018-0040-0001-0000
- Page Start:
- 94
- Page End:
- 101
- Publication Date:
- 2018-01
- Subjects:
- Transfer -- demonstration trajectories -- k-nearest neighbour -- reward shaping -- reinforcement learning
Automatic control -- Periodicals
Measuring instruments -- Periodicals
Commande automatique -- Périodiques
Mesure -- Instruments -- Périodiques
681.2 - Journal URLs:
- http://catalog.hathitrust.org/api/volumes/oclc/49488911.html ↗
http://tim.sagepub.com/ ↗
http://www.ingenta.com/journals/browse/arn/tm?mode=direct ↗
http://www.uk.sagepub.com/home.nav ↗ - DOI:
- 10.1177/0142331216649655 ↗
- Languages:
- English
- ISSNs:
- 0142-3312
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 8124.xml