DOI: 10.3390/pr14182995 ISSN: 2227-9717

RAPO-RL-TAC: Risk-Aware Partial-Order Reinforcement Learning with Timed Automata Completion for the Interval Job Shop Problem

Pujie Han, Yiheng Liu, Min Huang

In the Interval Job Shop Problem (IJSP), operation processing times are represented by intervals, making machine-sequencing decisions sensitive to temporal uncertainty. We propose Risk-Aware Partial-Order Reinforcement Learning with Timed Automata Completion (RAPO-RL-TAC), which couples Risk-Aware Partial-Order Reinforcement Learning (RAPO-RL) for machine-order construction with Timed Automata Completion (TAC) for execution-time resolution of deferred sequencing decisions. Risk awareness focuses on preserving temporal flexibility in machine-order relations whose preferred ordering is sensitive to interval uncertainty. RAPO-RL selectively commits comparatively determinate machine conflicts while retaining a bounded set of timing-sensitive relations. TAC completes unresolved relations as execution evolves, while statistical model checking characterises completion-time variability across timed executions. On 14 ORB and LA benchmark instances, RAPO-RL-TAC achieves mean midpoint makespans 2.92% and 1.63% lower than population-based neighbourhood search (PNS) and genetic algorithm (GA), respectively. Compared with reproduced Fast Elitist Artificial Bee Colony (fEABC) variants, RAPO-RL-TAC achieves a lower midpoint than at least one variant on 7 of 14 instances. In the controlled ablation study, RAPO-RL-TAC achieves a 13.85% lower mean midpoint makespan than hard enforcement of the learned relations. These results indicate that risk-aware selective commitment preserves temporal flexibility while maintaining competitive nominal schedule quality under interval uncertainty.