Real-Time Sound Event Localization and Detection for Edge and Mobile Devices: A Systematic Review
Phathorn Thammasorn, Porawat VisutsakBackground: Real-time Sound Event Localization and Detection (SELD) is important for enabling spatial-audio perception in resource-constrained environments. Edge and mobile platforms—including smartphones, wearables, embedded processors, microcontrollers, and IoT devices—impose strict constraints on latency, memory, and power that challenge current SELD methods. Objective: This review evaluates the readiness of current SELD methods for real-time deployment on edge and mobile devices and identifies the principal limitations affecting deployment. Methods: This systematic review was reported in accordance with the PRISMA 2020 statement. A structured search conducted on 30 June 2026 using SciSpace, SciSpace Full Text, Google Scholar, and arXiv retrieved 1078 records before deduplication, resulting in 611 unique records. Eligible studies were English-language publications from January 2018 to 30 June 2026 addressing SELD or directly relevant edge/mobile acoustic perception and deployment, with two foundational pre-2018 exceptions. Study selection and data extraction were performed by the first author, with automated tools used only to assist prioritization and information extraction. No formal risk-of-bias tool was applied. Because of methodological heterogeneity, results were synthesized qualitatively. Following screening and eligibility assessment, 49 publications were included. Results: Significant progress has been made in developing compact systems capable of performing SELD in real time; however, many limitations persist. Trade-offs between model complexity and localization accuracy remain unresolved. Sensitivity to microphone geometry and channel count is a persistent constraint. Robustness under noisy and reverberant conditions has improved, but current systems remain limited in practical deployment. A gap persists between performance on SELD benchmark datasets and real-world deployments. The available evidence was heterogeneous, and deployment-related metrics were reported inconsistently across studies. Conclusions: Several directions are identified that may help close these gaps, including hardware-aware Neural Architecture Search (NAS), self-supervised spatial-audio representation learning, on-device adaptation/personalization, energy-aware duty cycling, multimodal integration, cooperative device fusion, microphone calibration, and privacy-preserving trustworthy edge SELD. This systematic review provides an organized synthesis of the current state of the art in edge/mobile SELD, its limitations, and future opportunities for real-time implementation. The review was not prospectively registered and received no external funding.