DOI: 10.1371/journal.pone.0359641 ISSN: 1932-6203

Speech enhancement method integrating squeeze and excitation channel attention mechanism and improved CGAN

Yecai Guo, Meiyu Liang, Tianyue Jiang

To avoid the additional computational cost and algorithmic latency introduced by short-time Fourier transformation and waveform reconstruction, we propose CGAN-SECA, which incorporates a squeeze-and-excitation channel attention mechanism to enhance noisy speech directly in the time domain. The generator employs a symmetric encoder-decoder backbone with skip connections to reconstruct waveform samples and reduce information loss during down-sampling and up-sampling. SECA adaptively recalibrates channel-wise feature responses, enabling the network to emphasize speech-relevant representations under complex noise. In the discriminator, Virtual Batch Normalization (VBN) and Dropout are introduced to regularize adversarial training. Least-squares adversarial objectives replace cross-entropy-based objectives, and an L1 reconstruction term encourages sample-level fidelity. The method is evaluated on VoiceBank-DEMAND, which contains environmental and babble noise, and on the THCHS30-DNS Challenge dataset under multiple Signal-to-Noise Ratio (SNR) conditions. On VoiceBank-DEMAND, CGAN-SECA achieved PESQ and STOI scores of 3.018 and 0.954, respectively, ranking second only to CMGAN and outperforming the remaining comparison methods. CGAN-SECA also maintained fast inference, with a real-time factor (RTF) of 0.00423. On the THCHS30-DNS Challenge dataset, CGAN-SECA achieved the highest average PESQ and STOI scores among all evaluated methods, reaching 2.274 and 0.838, respectively, across the −5, 0, and 5 dB conditions. In particular, it produced the best results for both metrics at 0 and 5 dB, demonstrating effective speech enhancement and good generalization across different noise levels and datasets.