Single-point optical-vibration sensing system for deep-learning-based stereo sound synthesis
Kuo-Wei Chao, Ji-Yan Han, Jia-Wei Chen, Yi-Chieh Lin, Ying-Hui LaiStereo sound enhances the ability to identify the locations of sound sources, thereby enriching the auditory experience. The prevalent method for stereo recording involves using a microphone array, but this approach encounters challenges related to distance, ambient noise, and equipment setup complexity. To address these limitations, this paper proposes a deep-learning-based stereo sound synthesis system using a laser Doppler vibrometer. This study verifies that these monophonic vibration signals contain encoded binaural hearing cues, particularly the interaural level difference and interaural time difference. The proposed system leverages a gated convolutional neural network to synthesize stereo sound from these signals. The system achieves directional sound reconstruction from single-point vibration data with recognition rates of 97.7%–98.8%. The results of a stereo music synthesis experiment confirmed the effectiveness of the system, indicating scale-invariant signal-to-distortion ratio improvements of 2.64 and 3.52 for inside and outside tests, respectively. Subjective evaluations showed significant enhancements, with the mean opinion score for pleasantness and stereoscopic sense increasing by 44.5% and 57.0%, respectively.