DOI: 10.1002/aaai.70098 ISSN: 0738-4602

Designing healthy and resilient information environments: A multi‐agent sandbox for exploring risks and countermeasures

Jurriaan van Diggelen, Maaike Homan

Abstract

Artificial intelligence is transforming online information environments, amplifying both societal benefits and risks. This article argues that online platforms can be understood as multi‐agent systems (MAS) composed of humans, simulated users, adversarial agents, and defensive agents operating under platform‐level rules. In this MAS‐based framework, persuasion, trust, coordination, and governance interact to shape system‐level outcomes. To study these dynamics safely, we introduce a sandbox environment in which human participants, red bots, green bots, and blue bots can be observed under controlled conditions. Early experiences with the sandbox show that persuasive agents can be built with little effort, that current large language model (LLM) agents lack psychological realism, and that humans may struggle to distinguish malicious red bots from benign participants. The sandbox enables stakeholders to experience and evaluate trade‐offs in moderation, amplification, and intervention strategies. We outline a research agenda for improving agent validity, modelling advanced threats, and designing human–machine teams that support healthier, more resilient information environments.