Behavioral Analysis of Transfer Learning in DQN-Based Pedestrian Agents Using Single-Agent Pretraining
Tomohiro Hayashida, Shinya Sekizaki, Natsuki KatoReinforcement-learning-based pedestrian agents can acquire adaptive behaviors without hand-crafted motion rules, but training from scratch in each environment is computationally expensive and often environment dependent. This paper investigates a two-stage transfer-learning framework for Deep Q-Network (DQN)-based pedestrian agents and analyzes how single-agent pretraining changes subsequent learning and behavior formation in multi-agent environments. A shared Q-network is first pretrained over four single-agent source tasks using Conservative Q-Learning with adaptive source-task allocation and is then transferred as the initial network of every agent in four multi-agent target tasks. Unlike fixed curriculum training, the source tasks are not traversed in a predetermined order; instead, experience collection is reallocated according to current collision and goal-progress performance. Compared with random initialization, the transferred network improves early-stage learning in intersection, bi-flow, and bi-door environments and significantly reduces target-task training time in these three tasks. In the bottleneck environment, the improvement is limited and the reduction in target-task training time is not statistically significant. Behavioral analysis further shows that pretraining changes early exploration: in the bi-door task, transferred agents reach goal-side and door regions earlier while spending less time near walls. Post-training trajectories also reveal collision avoidance and diverse route choices. These results clarify both the benefits and limitations of single-agent pretraining for efficient and interpretable behavior acquisition in multi-agent pedestrian simulation.