How large language models judge and influence human cooperation
Alexandre S Pires, Laurens Samson, Sennay Ghebreab, Fernando P SantosAbstract
Humans increasingly rely on large language models (LLMs) to support decisions in social settings. Previous work suggests that such tools can shape people’s moral judgments. However, how LLMs evaluate decisions in social dilemmas, and the long-term implications of LLM-based assessments on human cooperation—driven by indirect reciprocity, reputations, and the capacity to judge interactions of others—remain unclear. We evaluate the social norms of 21 state-of-the-art LLMs when judging more than 40,000 fictitious decisions in social dilemmas. Through an evolutionary game-theoretical model, we study prosociality in populations adhering to these norms, theoretically studying their potential impact on long-term human cooperation. We observe a remarkable agreement in LLMs when evaluating actions against opponents perceived as good. Yet, there is large within- and between-model variance when assessing actions toward ill-reputed individuals. Our theoretical analysis reveals that these differences can significantly impact the prevalence of cooperation within our model. Finally, we show that prompt-based interventions can steer LLM norms, particularly when these define objective goals. Our theoretical model links LLM-based advice to long-term social dynamics, highlighting the importance of designing LLM norms considering their potential effects on human cooperation.