Comparative Analysis and Unified Evaluation Framework of User Simulators for Realistic Behaviour Modelling
Md Faisal Ahmed, Estrid He, Chenglong Ma, Jeffrey ChanHuman-in-the-loop experiments have been increasingly adopted in AI research. Nevertheless, they usually require elevated time/monetary costs and ethical considerations. Hence, developing user behaviour simulators to synthesise reliable and realistic human behaviours has emerged as an attractive alternative. Although various simulators have been proposed, there lacks a comprehensive experimental study that systematically benchmarks these simulators under a unified evaluation framework. In this paper, we fill this gap by comparing two mainstream simulators experimentally and comprehensively: reinforcement learning (RL) environments and large language models (LLMs) based, focusing on how they respond to different data settings, e.g., data sparsity and application domains. Our extensive experimental results reveal several key findings. First, LLM-based simulators demonstrate stronger performance in general, highlighting their significant potential to serve as an underpinning component of future works. However, they are found to be more sensitive to data density and biased towards positive interactions: their performance on sparse and negative bias conditions on datasets like Book Crossing drops significantly by 44% to 58% despite their strong performances on denser datasets. RL environment simulators are more effective in capturing users’ long-term preference profiles, evidenced by their capability in modelling the user conformity phenomenon over time. Finally, LLM-based simulators tend to overemphasise positive feedback, often overlooking negative signals compared to RL environment simulators.