DOI: 10.3390/systems14091172 ISSN: 2079-8954

Resilience Metric and Enhancement for Adversarial Operational Loop Networks via Reinforcement Learning

Jidong Sun, Xiaobo Li, Zhi Zhu, Wei Li, Tao Wang, Zhijie Huang, Jie Zhang

Resilience optimization in adversarial Multi-Agent Mission Systems (MAMSs) requires a decision mechanism whose objective is explicitly aligned with the resilience metric, not merely resilience measurement. This paper proposes a net-effectiveness-based evaluation–optimization framework for bilateral temporal adversarial operational-loop networks (BTAR), turning resilience evaluation into a reusable world model and reward source for decision-making. We model bilateral confrontation as a heterogeneous temporal network and quantify resilience through the integral of net effectiveness, which captures temporal and adversarial dynamics beyond static topological indicators. We then formulate resilience enhancement as a fixed-dimensional, size-agnostic, cost-aware Markov decision process and solve it with a hierarchically decoupled reinforcement learning architecture: an upper-level PPO policy learns when and for which chains to trigger selective replanning, while a lower-level greedy routine selects concrete replacement nodes. In this way, the optimization objective remains aligned with the evaluation metric through a metric-aligned reward. Experiments on a reproducible multi-scale, multi-style scenario pool show that the learned policy improves Rtotal from 0.3673 to 1.1696, reaching 96.8% of an idealized oracle benchmark (1.2082), while slightly outperforming the always-replan rule (1.1482) with a lower trigger rate. Under nonzero replanning cost, retraining yields a cost-aware policy whose trigger rate drops to 21.2% and whose cost-adjusted resilience at the target cost setting increases by +0.15, surpassing the greedy oracle benchmark in the high-cost regime. These results show that coupling net-effectiveness evaluation with hierarchical RL yields a deployable, size-agnostic, and cost-aware resilience optimizer.