DOI: 10.3390/computation14090223 ISSN: 2079-3197

Solver-Independent MAPPO for Disruption-Aware Last-Mile Routing via Candidate Parcel-Locker Consolidation Points

Mohamed-Ali Ejjanfi, Kadim Lahcen Nadime, Jamal Benhra

Urban last-mile routes must adapt when demand and network conditions change after dispatch. This study proposes a solver-independent multi-agent proximal policy optimization (MAPPO) framework that combines a shared fleet policy, dynamically ranked candidate parcel-locker consolidation points, disruption-aware observations, fleet-level value learning, and structural feasibility masking. Evaluation on a Los Angeles road-network simulation used chronological data splits, ten random seeds, matched routing baselines, and compound-disruption stress tests. MAPPO served 98.4% of demand versus 96.6% for static OR-Tools and reduced failed packages by 52.2%; under severe compound disruption, service increased from 89.4% to 93.0%. Higher-service adaptive methods reached 99.6–99.7%, with summed online computation times of 0.111–0.503 s per episode compared with 0.029 s for MAPPO. The integrated framework therefore improves disruption recovery over static planning while exposing a clear quality–computation trade-off against adaptive search.