Exploration by Multi-Agent Relationship Graph Reconstruction
Yaxin Xu, Yinxiang He, Tianyi Liu, Ningzhong Liu, Han Sun, Mingjing Pei, Jiaquan ShenEfficient exploration remains a key challenge in cooperative multi-agent reinforcement learning (MARL), where the novelty of a joint situation may arise not only from unfamiliar individual observations but also from previously underrepresented interaction patterns among agents. Existing intrinsic-reward approaches commonly characterize novelty through state, observation, value, or prediction representations, which may not explicitly capture such relational changes. To address this issue, we propose Relationship Graph Reconstruction Exploration (RGRE), an intrinsic exploration method that characterizes novelty from the perspective of inter-agent relationships. RGRE first constructs an attention-derived relationship representation from agents’ local observations and then employs a graph autoencoder to learn its latent relational structure. The reconstruction discrepancy is used as a proxy for relational novelty and incorporated into the environmental reward as an intrinsic exploration bonus. Experiments on the Multi-Agent Particle Environment (MPE) Spread task with 3, 6, and 10 agents and on six StarCraft Multi-Agent Challenge (SMAC) scenarios show that RGRE achieves higher mean final performance than QMIX on all nine evaluated tasks. Based on the final-performance results, the average improvements are 5.82 percentage points in landmark occupation rate across the three MPE settings and 7.84 percentage points in test win rate across the six SMAC scenarios.