DOI: 10.1145/3832005 ISSN: 2474-9567

ASAP: Acoustic-Semantic Alignment with Prototypes for Open-world Activity Recognition and Understanding on Glasses

Changfei Dong, Qian Zhang, Dong Wang

Acoustic sensing-enabled smart glasses offer a privacy-preserving substrate for continuous activity understanding by probing near-body motion with inaudible acoustics. However, existing solutions are typically trained as closed-set classifiers: expanding the activity vocabulary requires costly data collection, and accuracy degrades sharply under cross-user and cross-environment shift. We present ASAP (Acoustic-Semantic Alignment with Prototypes), an open-world activity recognition system for ultrasonic smart glasses that targets practical near-neighbor vocabulary growth and robust generalization in daily life. ASAP converts ultrasonic echoes into motion-sensitive features and aligns them with text-derived activity prototypes in a shared acoustic-language embedding space. At inference time, ASAP integrates similarity-based rejection with two-stage seen/unseen routing, and supports rapid personalization via few-shot prototype fusion without retraining. Across 26 participants with 20 seen and 7 unseen activities, ASAP achieves 71.10% harmonic-mean accuracy under generalized zero-shot recognition, improves to 80.23% with few-shot fusion at k =3, and boosts cross-user generalization over a strong baseline by up to 18.8 points. For long-form continuous streams, ASAP further leverages an LLM as a post-hoc calibrator and journal generator, raising the harmonic mean to 80.87% (zero-shot) and 88.73% (few-shot) on in-the-wild long sessions, while user feedback shows that generated journals are preferred over raw recognition streams.