Speech enhancement method integrating squeeze and excitation channel attention mechanism and improved CGAN
Yecai Guo, Meiyu Liang, Tianyue JiangTo avoid the additional computational cost and algorithmic latency introduced by short-time Fourier transformation and waveform reconstruction, we propose CGAN-SECA, which incorporates a squeeze-and-excitation channel attention mechanism to enhance noisy speech directly in the time domain. The generator employs a symmetric encoder-decoder backbone with skip connections to reconstruct waveform samples and reduce information loss during down-sampling and up-sampling. SECA adaptively recalibrates channel-wise feature responses, enabling the network to emphasize speech-relevant representations under complex noise. In the discriminator, Virtual Batch Normalization (VBN) and Dropout are introduced to regularize adversarial training. Least-squares adversarial objectives replace cross-entropy-based objectives, and an L1 reconstruction term encourages sample-level fidelity. The method is evaluated on VoiceBank-DEMAND, which contains environmental and babble noise, and on the THCHS30-DNS Challenge dataset under multiple Signal-to-Noise Ratio (SNR) conditions. On VoiceBank-DEMAND, CGAN-SECA achieved PESQ and STOI scores of 3.018 and 0.954, respectively, ranking second only to CMGAN and outperforming the remaining comparison methods. CGAN-SECA also maintained fast inference, with a real-time factor (RTF) of 0.00423. On the THCHS30-DNS Challenge dataset, CGAN-SECA achieved the highest average PESQ and STOI scores among all evaluated methods, reaching 2.274 and 0.838, respectively, across the −5, 0, and 5 dB conditions. In particular, it produced the best results for both metrics at 0 and 5 dB, demonstrating effective speech enhancement and good generalization across different noise levels and datasets.