POPKI: Modeling Surprise in a Physically-Grounded Joint Inference of Observation, Preference, and Knowledge
Harry Chen, Karen Chung, Abhishek Bhandwaldar, Joshua B. Tenenbaum, Tomer D. Ullman, Tianmin ShuAbstract
People naturally take into account others’ limited perceptual access when attributing mental states such as preferences and knowledge. This ability is central to intuitive psychology, and even infants recognize when others have only partial observability. Understanding how mental state attributions adapt under limited perceptual access is thus crucial for advancing models of intuitive psychology. Here, we introduce POPKI (Physically-grounded Observation, Preference, and Knowledge Inference), a Bayesian inverse-planning framework designed to model graded surprise based on agents’ inferred preferences and knowledge in visually constrained 3D environments. Building on and expanding the AGENT dataset, we develop new scenarios that test how partial observability shapes judgments of surprise. Despite the tremendous progress of state-of-the-art vision-language models (VLMs), our experiments show that they fail to capture the nuances of perceptual access when judging surprise. POPKI, in contrast, offers a principled approach to modeling human-like uncertainty under partial observability and closely replicates human surprise patterns. It provides a foundation for advancing psychological reasoning models and enhancing human-AI interaction.