DOI: 10.3390/electronics15153369 ISSN: 2079-9292

HC-MARL: A General Hierarchical Cascaded Multi-Agent Collaborative Architecture for Cyber Wargaming Applications

Zhiqiang Qu, Jun He, Bo Wu, Zhitao Long, Tao Xia

Cyber wargaming serves as a core tool for simulating cyber confrontations and supporting operational decision verification. Existing multi-agent methods face issues such as the absence of cross-level mechanisms and poor adaptability to dynamic environments when applied to cyber wargaming. To address these challenges, we innovatively propose HC-MARL, a general hierarchical cascaded multi-agent reinforcement learning architecture tailored for cyber wargaming. Within this architecture, agents are modeled as hierarchical cascaded units to achieve structural decoupling and preserve scalability in both horizontal and vertical dimensions. Specifically, we devise a cross-level bidirectional information passing and intrusion alert sharing mechanism to accommodate cross-level command characteristics; design a Transformer-based message transformation function that fuses variable-length observation vectors from lower-level agents into fixed-length vectors, adapting to dynamic changes in agent structures and numbers; and design a policy function that integrates neural networks with empirical knowledge, incorporating an intrusion-alert-based action mask to enhance threat response efficiency. To the best of our knowledge, the proposed HC-MARL framework is a novel general hierarchical collaborative multi-agent architecture for cyber wargaming. To validate the effectiveness of our framework, we instantiate two-layer and three-layer agent clusters and conduct experiments in the CybORG CC4 environment. The results demonstrate that our architecture can effectively accommodate cross-level information transfer and dynamic changes in agents in cyber wargaming scenarios. Furthermore, incorporating intrusion-alert-based action masking into the policy function significantly improves defensive performance. Compared with the state-of-the-art (SOTA) methods, the average reward improves by approximately 12.3%, and compared with mainstream methods such as Singh et al., the average reward improves by approximately 27.5%.

More from our Archive