Combination of learning from non-optimal demonstrations and feedbacks using inverse reinforcement learning and Bayesian policy improvement. (1st December 2018)
- Record Type:
- Journal Article
- Title:
- Combination of learning from non-optimal demonstrations and feedbacks using inverse reinforcement learning and Bayesian policy improvement. (1st December 2018)
- Main Title:
- Combination of learning from non-optimal demonstrations and feedbacks using inverse reinforcement learning and Bayesian policy improvement
- Authors:
- Ezzeddine, Ali
Mourad, Nafee
Araabi, Babak Nadjar
Ahmadabadi, Majid Nili - Abstract:
- Highlights: Learning from non-optimal demonstrations in presence of human evaluative feedbacks. Extending inverse reinforcement learning algorithm to incorporate human feedbacks. Our approach overcomes the challenge of non-optimality in demonstrations. Good performance of our approach in both simulated and real-world experiments. Abstract: Inverse reinforcement learning ( IRL ) is a powerful tool for teaching by demonstrations, provided that sufficiently diverse and optimal demonstrations are given, and learner agent correctly perceives those demonstrations. These conditions are hard to meet in practice; as a trainer cannot cover all possibilities by demonstrations, he may partially fail to follow the optimal behavior. Also, trainer and learner have different perceptions of the environment including trainer's actions. A practical way to overcome these problems is using a combination of trainer's demonstrations and feedbacks. We propose an interactive learning approach to overcome the challenge of non-optimal demonstrations by integrating human evaluative feedbacks with the IRL process, given sufficiently diverse demonstrations and the domain transition model. To this end, we develop a probabilistic model of human feedbacks and iteratively improve the agent policy using Bayes rule. We then integrate this information in an extended IRL algorithm to enhance the learned reward function. We examine the developed approach in one experimental and two simulated tasks; i.e., a gridHighlights: Learning from non-optimal demonstrations in presence of human evaluative feedbacks. Extending inverse reinforcement learning algorithm to incorporate human feedbacks. Our approach overcomes the challenge of non-optimality in demonstrations. Good performance of our approach in both simulated and real-world experiments. Abstract: Inverse reinforcement learning ( IRL ) is a powerful tool for teaching by demonstrations, provided that sufficiently diverse and optimal demonstrations are given, and learner agent correctly perceives those demonstrations. These conditions are hard to meet in practice; as a trainer cannot cover all possibilities by demonstrations, he may partially fail to follow the optimal behavior. Also, trainer and learner have different perceptions of the environment including trainer's actions. A practical way to overcome these problems is using a combination of trainer's demonstrations and feedbacks. We propose an interactive learning approach to overcome the challenge of non-optimal demonstrations by integrating human evaluative feedbacks with the IRL process, given sufficiently diverse demonstrations and the domain transition model. To this end, we develop a probabilistic model of human feedbacks and iteratively improve the agent policy using Bayes rule. We then integrate this information in an extended IRL algorithm to enhance the learned reward function. We examine the developed approach in one experimental and two simulated tasks; i.e., a grid world navigation, a highway car driving system and a navigation task by the e-puck robot. Obtained results show significant improved efficiency of the proposed approach in face of having different levels of non-optimality in demonstrations and the number of evaluative feedbacks. Graphical abstract: … (more)
- Is Part Of:
- Expert systems with applications. Volume 112(2018)
- Journal:
- Expert systems with applications
- Issue:
- Volume 112(2018)
- Issue Display:
- Volume 112, Issue 2018 (2018)
- Year:
- 2018
- Volume:
- 112
- Issue:
- 2018
- Issue Sort Value:
- 2018-0112-2018-0000
- Page Start:
- 331
- Page End:
- 341
- Publication Date:
- 2018-12-01
- Subjects:
- Teaching by demonstrations -- Inverse reinforcement learning -- Interactive learning -- Human evaluative feedbacks
Expert systems (Computer science) -- Periodicals
Systèmes experts (Informatique) -- Périodiques
Electronic journals
006.33 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09574174 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.eswa.2018.06.035 ↗
- Languages:
- English
- ISSNs:
- 0957-4174
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3842.004220
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 7159.xml