SoundTrace: Integrating Temporal Context and Episodic Memory for Real-Time Sound Recognition
Dhruv Jain, Jason MillerEnvironmental sound recognition systems have become increasingly capable, yet they often operate in context-free modes that ignore temporal continuity, environmental patterns, and user feedback. We present SoundTrace, a real-time sound recognition system that integrates temporal context and episodic memory to support adaptive, interpretable inference in everyday environments. SoundTrace augments a neural audio classifier with lightweight memory structures that store symbolic event traces, estimate scene-time and short-range sequence priors, and update those priors through user feedback. During inference, these memory-derived priors are retrieved and fused with model predictions to stabilize labels and provide interpretable reasoning. In a controlled evaluation, we show that contextual inference improves accuracy, reduces label volatility, and enhances robustness under ambiguous conditions. We also report findings from an eight-week in-home deployment with 14 deaf and hard-of-hearing participants, revealing how context-aware feedback and explanations shape users' trust, understanding, and correction strategies. Our results demonstrate the viability of context-integrated sound recognition in everyday environments.