A multi-action deep reinforcement learning framework for flexible Job-shop scheduling problem. (1st November 2022)
- Record Type:
- Journal Article
- Title:
- A multi-action deep reinforcement learning framework for flexible Job-shop scheduling problem. (1st November 2022)
- Main Title:
- A multi-action deep reinforcement learning framework for flexible Job-shop scheduling problem
- Authors:
- Lei, Kun
Guo, Peng
Zhao, Wenchao
Wang, Yi
Qian, Linmao
Meng, Xiangyin
Tang, Liansheng - Abstract:
- Highlights: An end-to-end DRL-based framework is introduced to solve the FJSP. Multi-PPO is used to learn job operation action and machine action sub-policies in MPGN. The proposed DRL shows its robustness via random and benchmark test instances. Abstract: This paper presents an end-to-end deep reinforcement framework to automatically learn a policy for solving a flexible Job-shop scheduling problem (FJSP) using a graph neural network. In the FJSP environment, the reinforcement agent needs to schedule an operation belonging to a job on an eligible machine among a set of compatible machines at each timestep. This means that an agent needs to control multiple actions simultaneously. Such a problem with multi-actions is formulated as a multiple Markov decision process (MMDP). For solving the MMDPs, we propose a multi-pointer graph networks (MPGN) architecture and a training algorithm called multi-Proximal Policy Optimization (multi-PPO) to learn two sub-policies, including a job operation action policy and a machine action policy to assign a job operation to a machine. The MPGN architecture consists of two encoder-decoder components, which define the job operation action policy and the machine action policy for predicting probability distributions over different operations and machines, respectively. We introduce a disjunctive graph representation of FJSP and use a graph neural network to embed the local state encountered during scheduling. The computational experiment resultsHighlights: An end-to-end DRL-based framework is introduced to solve the FJSP. Multi-PPO is used to learn job operation action and machine action sub-policies in MPGN. The proposed DRL shows its robustness via random and benchmark test instances. Abstract: This paper presents an end-to-end deep reinforcement framework to automatically learn a policy for solving a flexible Job-shop scheduling problem (FJSP) using a graph neural network. In the FJSP environment, the reinforcement agent needs to schedule an operation belonging to a job on an eligible machine among a set of compatible machines at each timestep. This means that an agent needs to control multiple actions simultaneously. Such a problem with multi-actions is formulated as a multiple Markov decision process (MMDP). For solving the MMDPs, we propose a multi-pointer graph networks (MPGN) architecture and a training algorithm called multi-Proximal Policy Optimization (multi-PPO) to learn two sub-policies, including a job operation action policy and a machine action policy to assign a job operation to a machine. The MPGN architecture consists of two encoder-decoder components, which define the job operation action policy and the machine action policy for predicting probability distributions over different operations and machines, respectively. We introduce a disjunctive graph representation of FJSP and use a graph neural network to embed the local state encountered during scheduling. The computational experiment results show that the agent can learn a high-quality dispatching policy and outperforms handcrafted heuristic dispatching rules in solution quality and meta -heuristic algorithm in running time. Moreover, the results achieved on random and benchmark instances demonstrate that the learned policies have a good generalization performance on real-world instances and significantly larger scale instances with up to 2000 operations. … (more)
- Is Part Of:
- Expert systems with applications. Volume 205(2022)
- Journal:
- Expert systems with applications
- Issue:
- Volume 205(2022)
- Issue Display:
- Volume 205, Issue 2022 (2022)
- Year:
- 2022
- Volume:
- 205
- Issue:
- 2022
- Issue Sort Value:
- 2022-0205-2022-0000
- Page Start:
- Page End:
- Publication Date:
- 2022-11-01
- Subjects:
- Flexible job-shop scheduling problem -- Multi-action deep reinforcement learning -- Graph neural network -- Markov decision process -- Multi-proximal policy optimization
Expert systems (Computer science) -- Periodicals
Systèmes experts (Informatique) -- Périodiques
Electronic journals
006.33 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09574174 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.eswa.2022.117796 ↗
- Languages:
- English
- ISSNs:
- 0957-4174
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3842.004220
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 21853.xml