DOI: 10.1145/3831648 ISSN: 2474-9567

MutterMeter: Earable Sensing of Self-Talk through Hierarchical Acoustic-Linguistic Fusion toward Self-Talk Interventions

Euihyeok Lee, Seonghyeon Kim, SangHun Im, Heung-Seon Oh, Seungwoo Kang

Self-talk is a meaningful psychological signal that reflects individuals' emotions, motivations, and cognitive states, yet remains difficult to capture in daily life. It often emerges spontaneously during moments of pressure, concentration, or repeated decision-making, making it a relevant target for behavioral and cognitive interventions such as Educational Self-Talk Intervention (ESTI). However, ESTI largely depends on expert observation or retrospective self-reports, which are limited in capturing fleeting self-talk patterns and hinder its scalability in real-world settings. To address this gap, we present MutterMeter, a mobile self-talk detection system that continuously analyzes audio from earable microphones and classifies utterances into Negative self-talk, Positive self-talk, and Others. Focusing on tennis as a high-pressure and cognitively demanding testbed, MutterMeter employs a hierarchical framework that integrates acoustic, linguistic, and contextual information while adaptively balancing accuracy and efficiency. Evaluated on a first-of-its-kind dataset we collected (34.5h, n=29), MutterMeter achieves a macro-averaged F 1 score of 0.815 across subjects, outperforming baselines such as LLM-based models and speech emotion recognition models. Moreover, its adaptive processing pipeline reduces unnecessary computation by finalizing confident cases early, while maintaining robust detection performance.