DOI: 10.3390/pr14152519 ISSN: 2227-9717

Adaptive Lagrangian Penalty-Enhanced Proximal Policy Optimization for Flexible Job Shop Rescheduling with Worker Workload Constraints Under Concurrent Dynamic Disturbances

Yuanmeng Zhou, Haoyi Tan, Jiawei Li

When flexible job shop scheduling faces concurrent disturbances such as machine failures and rush orders, worker-centric constraints emphasized under Industry 5.0 must also be satisfied. Existing deep reinforcement learning methods for the Dynamic Flexible Job Shop Scheduling Problem (DFJSP) seldom treat worker workload balance as an explicit constraint, and most depend on static penalty coefficients that are difficult to tune across different scenarios. In this paper, we suggest ALP-PPO, an adaptive Lagrangian penalty-enhanced proximal policy optimization algorithm, for real-time rescheduling under concurrent machine breakdowns and rush orders. We formulate the scheduling environment as a constrained Markov decision process. Worker skill heterogeneity, fatigue accumulation and workload equity are modeled as coupled constraints alongside classical scheduling objectives. By decoupling operation sequencing, machine allocation and worker assignment into coordinated sub-decisions, a hierarchical action space is constructed. Dual Lagrangian multipliers for workload balance and fatigue are updated adaptively during training, so that manual penalty tuning is no longer required. An event-triggered mechanism selects between right-shift and full rescheduling on the basis of a disruption severity index. We employ weighted-sum scalarization of makespan, energy consumption and workload variance during training, and Pareto solution sets are obtained by systematically varying the weight vectors across independent training runs. On extended Brandimarte benchmarks augmented with worker and dynamic event parameters, ALP-PPO delivers superior scheduling performance across makespan, energy consumption and workload variance when compared with Double DQN, Dueling DQN, standard PPO, NSGA-II and MOEA/D, as measured by Hypervolume (HV) and Inverted Generational Distance (IGD) indicators. Ablation studies indicate that the adaptive Lagrangian mechanism reduces constraint violations by more than 40% relative to fixed-penalty alternatives while keeping the primary objectives competitive. An analysis of computational efficiency shows that ALP-PPO completes online inference in under 20 ms per decision step, making real-time rescheduling practically feasible. Generalization experiments on previously unseen instances further validate the transferability of the learned policy. These findings support human-centric intelligent scheduling in Industry 5.0 manufacturing.

More from our Archive