DOI: 10.1145/3831653 ISSN: 2474-9567

Feasibility of Using a Multi-Agent LLM System to Correct Annotations and Support Low-Effort Activity Labeling

Ha Le, Akshat Choube, Varun Mishra, Stephen S. Intille

Accurate measurement of behaviors is critical for research in human-computer interaction, ubiquitous computing, and personal health informatics, because it underpins many health tracking and intervention systems. Most data collection studies, however, still rely on participants' self-reports and manual annotations, which often under- or over-estimate activity duration and type, and require substantial effort from researchers to clean and validate. An automated system that can combine passive sensing data with participants' self-reports might detect inconsistencies and suggest corrections. We introduce GLOSS4HAR, a multi-agent LLM-based system designed to mimic human sensemaking and assist researchers in cleaning and refining activity annotations. We demonstrate the potential of GLOSS4HAR in two key tasks: (1) correcting and reconciling participant self-annotations, and (2) triangulating passive sensing data with different forms of lightweight self-reports to generate accurate activity timelines. Our evaluation shows that GLOSS4HAR improves annotation quality by up to 9.9% in F1 score and can reconstruct activity timelines that align with human annotations at 75-92% F1. Based on our findings, we discuss the implications of our work for the next generation of activity annotation systems that might use human-AI collaboration.