Reinforcement-learning-based replicator–mutator dynamics for game-environment coevolution in public goods games
Xinyu Liu, Linchao Pan, Chunying Ren, Bangxin Jiang, Changbing TangThe reciprocal feedback between individual strategies and environmental states has received widespread attention in mathematical biology and ecology. However, existing game-environment feedback models typically focus on myopic payoff-driven adaptation, overlooking how individuals adjust their behavior based on long-term learning outcomes. In this paper, we derive a learning-induced mutation term from frequency-adjusted Q-learning with Boltzmann exploration to describe how long-term value learning shapes exploratory strategy adjustment. Based on this, we then embed the learning-induced mutation term into game-environment feedback dynamics, leading to reinforcement-learning-based replicator–mutator dynamics for public goods games. In the above coupled framework, the cooperation frequency affects the environmental state, while the environmental state reshapes the payoff difference and, hence, the strategy dynamics. More importantly, the derived mutation term is nonlinear and frequency-dependent, capturing exploratory strategy adjustments based on long-term value learning. Specifically, the learning-rate scale plays a crucial role in system dynamics, and appropriate learning-rate scales can help maintain positive cooperation and mitigate the tragedy of the commons. In addition, a low learning-rate scale may weaken the local damping around the interior equilibrium and lead to long-lived damped strategy-environment oscillations with reduced amplitude, providing insights into how learning-induced exploration regulates coevolutionary fluctuations.