Reinforcement Learning-Based Day-to-Day Route Choice Models Under Full and Partial Information Conditions: A Theoretical and Experimental Research
Yuance Yang, Ning Jia, Nianlu Ren, Hongye Fan, Zhanghao LeiLearning behavior plays an important role in the day-to-day route choice process, which is essentially a repeated decision-making process. This paper proposes two learning models for the day-to-day route choice process, one for the full-information (FI) condition and one for the partial-information (PI) condition. Both models are developed based on attraction-based reinforcement learning theory but with different learning mechanisms. In our FI model, travelers learn by comparing different paths, while in our PI model, travelers’ route choice decision is modeled by long-term cost minimization and a policy-based reinforcement algorithm. Both models are rooted in individual strategy updating and are theoretically linked to classical network flow assignment theory: we establish that individual-level invariance of choice probabilities is sufficient for aggregate stationarity consistent with Wardrop equilibrium, and provide an explicit stability condition for the FI dynamics. To validate the models, laboratory experiments under FI and PI conditions were designed and conducted. The experimental data confirm the presence of learning behavior and show that the proposed models achieve better predictive performance than several behavioral benchmarks, including EWA, Q-learning, mixed logit, and a Selten variant. Moreover, an interesting phenomenon was found in the FI experimental dataset: even though FI was provided, some subjects still made decisions hinging on their own experience. Our findings would benefit traffic information services and traffic administration.