Reinforcement learning with reputation-based adaptive exploration promotes cooperation
An Li, Wenqiang Zhu, Chaoqian Wang, Longzhao Liu, Hongwei Zheng, Yishen Jiang, Xin Wang, Shaoting TangReinforcement learning provides a framework for studying how individuals adjust their behavior through repeated interaction and feedback in social dilemmas. In Q-learning, exploration controls how often agents choose actions other than those favored by their current learned Q-values. Yet, the existing models usually treat the exploration rate as a constant parameter. In systems with social evaluation, however, trial-and-error behavior carries different costs and opportunities for agents with different reputations, making exploration dependent on social standing rather than uniform across agents. Herein, we develop a spatial prisoner’s dilemma model in which Q-learning agents adapt their exploration rates according to local reputation differences, while reputation is updated through an asymmetric, state-dependent rule. The results show that adaptive exploration and asymmetric reputation updating each promote cooperation, but their combination produces a stronger reinforcing effect than either mechanism alone. Low-reputation agents explore more and can recover reputation through cooperation, while high-reputation agents explore less and avoid reputation losses caused by defection. This mechanism also reorganizes cooperation in space, producing a stable checkerboard-like coexistence at intermediate reputation concern. In addition, cooperation is most vulnerable at intermediate baseline exploration rates, whereas stronger asymmetric reputation updating mitigates this exploration-induced disruption. These results suggest that reputation can act not only as a record of past behavior but also as a dynamic signal that regulates exploratory behavior during learning and thereby stabilizes cooperation.