MIRA: Safety-Constrained Multi-Agent Reinforcement Learning for Joint Prescriptive Maintenance and Production Rescheduling in Industrial IoT
Md. Ashraful Babu, Ali AlArjani, Mohamed LahbyIndustrial IoT maintenance often stops at health prediction, leaving maintenance, rescheduling, safety, and communication to separate decision processes. This study presents MIRA, a safety-constrained graph-based multi-agent reinforcement learning architecture for joint prescriptive maintenance, production rescheduling, and event-triggered communication. Machine condition was estimated from CNC milling data using temporal convolutional models; because predictive uncertainty failed a predefined validation gate, the controller used deterministic health estimates. Evaluation covered five controllers, six simulated scenarios, and 1800 matched episodes. Relative to Graph-MAPPO, MIRA reduced operational cost by 9.38%, weighted tardiness by 28.10%, unexpected failures by 17.39%, message count by 84.98%, and transmitted data by 83.83%, while increasing on-time completion by 23.55%, without a detectable difference in corrected critical-message recall. Across the three independently trained seeds, failures, safety violations, and message count favored MIRA consistently, whereas cost and tardiness favored MIRA in two seeds. Disabling the execution shield increased safety violations from 0 to 3.56 per episode. Post-training variation in the projected-health safe-start threshold from 0.124 to 0.132 produced no safety violations and only small changes in aggregate operational outcomes. Cross-domain health transfer to PHM 2010 failed without adaptation. The results support simulator-level decision coordination, while broader replication, variable-size deployment, and factory validation remain necessary.