DOI: 10.3390/electronics15153405 ISSN: 2079-9292

Decentralized Learning and Control of Multi-Microrobots in Complex Hemodynamic Environments

Truong Nhut Huynh, Kim-Doang Nguyen

Autonomous microrobot teams have significant potential for distributed drug delivery, cooperative vascular intervention, and parallelized biomedical diagnostics. However, coordinated control in cardiovascular environments remains challenging due to partial observability, limited communication bandwidth, and hydrodynamically coupled pulsatile blood flow. This paper introduces Decentralized Hemodynamic-Aware Multi-Agent Reinforcement Learning (DH-MARL), a distributed learning and control framework in which individual microrobots learn decentralized policies from local observations while graph-based attention mechanisms model inter-agent interactions during centralized training. The proposed framework integrates turbulence-aware adaptive exploration, reduced-order hydrodynamic interaction modeling, diffusion-based local communication, and hemodynamic-aware counterfactual credit assignment to improve cooperative learning and role specialization in dynamic vascular environments. A scalable Unity-based simulator supporting coupled pulsatile flow for up to 32 agents was developed for training and evaluation. Our experimentalresults cover four therapeutic scenarios: distributed drug delivery, cooperative clot dispersion, stenosis mapping, and vessel bottleneck traversal. For 16-agent teams, DH-MARL reaches an 88.7% team success rate. This performance exceeds independent single-agent controllers and centralized MAPPO baselines, and inter-robot collision rates remain below 4%. The learned policies generalize to unseen team sizes with minimal performance degradation, highlighting the scalability and robustness of the proposed decentralized control strategy. These simulation-level results demonstrate the feasibility of distributed reinforcement learning and graph-based coordination as a control paradigm for future multi-agent microrobot systems in biomedical environments. The results also provide a foundation for the calibration of subsequent microfluidic, ex vivo, and preclinical experiments.

More from our Archive