Auditing Single-Agent Reinforcement Learning for EV Charging Assignment: A Protocol-Amended Comparison of Trained, Untrained, and Heuristic Policies
Nour-Eddine Moumni, Rachid Alaoui, Driss KiouachDeep reinforcement learning (DRL) is widely assumed to outperform simpler rule-based and tabular baselines for sequential decision problems. We test this assumption for electric vehicle (EV) charging assignment using Simulation of Urban MObility (SUMO) simulations of real Rabat and Tangier road networks, with a protocol-amended, trained-vs-untrained diagnostic of Deep Q-Network (DQN) and Double DQN (DDQN) agents against untrained controls and an adaptive heuristic under a corrected pipeline and a not-yet-field-validated station-power regime. The primary endpoint, mean queue waiting time, is an assignment-to-arrival access delay. Trained policies are classified against a pre-specified 60 s rule using multiplicity-adjusted bootstrap intervals: of 12 pre-defined contrasts, none reaches the rule, two show a small detectable advantage, one is detrimental (softening to inconclusive with ten seeds), and nine are inconclusive. A legacy 210-run campaign predating the pipeline corrections is exploratory and shows no Holm-corrected difference. A reward/Markov decision process (MDP) audit identifies four structural defects and one replay sampling limitation; of the five audit-derived interventions, three were already active, one was structurally inert although enabled, and one showed no improvement. Exact per-contrast tests of the isolated ablations cannot reject at seven seeds; restoring R2 yields a small DDQN effect in one scenario that does not reproduce in a second. Two post hoc exploratory comparators, a deterministic minimum expected completion time dispatcher (MECT) and a discrete action proximal policy optimization (PPO) agent, are reported descriptively. The diagnostic does not separate under-training from a non-discriminative environment. The broader contribution is a reusable audit workflow for applied reinforcement learning benchmarks: endpoint semantic validation, trained-vs-untrained controls, multiplicity-adjusted inference and reward specification inspection.