Dual-path magnitude-phase learning with bidirectional cross-attention for bone-conducted speech restoration
Dianzhe Ding, Sichen Liu, Junfeng Li, Feiran YangBone-conducted speech is not subject to background noise but suffers from limited bandwidth. In very low signal-to-noise ratio conditions, it is highly desired to restore air-conducted speech from bone-conducted speech. This paper presents a dual-path frequency-domain U-shaped Convolutional Network (UNet) model to convert bone-conducted speech into high-quality air-conducted speech, which jointly utilizes the magnitude and phase spectrograms of bone-conducted speech. Compared with the widely adopted single-path model for bone-conducted speech restoration and dual-path model for speech denoising, the proposed model consists of dual-path convolution modules in the encoder and decoder and a bidirectional cross-attention module in the bottleneck layer. Dual-path convolution modules utilize two individual branches to extract the magnitude and phase features from the bone-conducted speech spectrogram, which aims to fully exploit the information intrinsic to each spectrogram. Because the two features provide distinct information regarding bone-conducted speech, a bidirectional cross-attention module is introduced to help the model learn consistent information from both features. Moreover, we propose a frequency-weighted spectrogram loss as the training objective. Experiments show that the proposed model outperforms existing bone-conducted speech restoration models, while utilizing only 7.6 M parameters and 2.35 G multiply-accumulate operations, which is lower than that of most competitive methods.