Machine Learning-Enhanced Uniform DFT Polyphase Filter Bank for Real-Time Speech Enhancement in Health Care Applications
Thaseena C.K., Gurinder SodhiPurpose – To improve speech quality and intelligibility in noisy healthcare environments, such as ICUs, emergency wards, and telemedicine systems, using a machine learning-enhanced speech enhancement framework. Design/methodology/approach – The proposed method uses a Uniform DFT Polyphase Filter Bank (PFB) to divide noisy speech into frequency subbands. Machine learning models, including CNNs, LSTMs, and Transformers, estimate noise masks and enhance speech before reconstructing the clean signal. Findings – Results show significant improvements in speech quality and intelligibility, with SNR gains of 5–12 dB and higher PESQ and STOI scores compared to traditional methods. CNN and Transformer models performed particularly well in complex hospital noise conditions while maintaining low latency for real-time use. Originality/value – The study combines efficient filter bank processing with advanced machine learning techniques to provide a practical and scalable solution for real-time speech enhancement in healthcare communication systems.