Psychoacoustically inspired optimization framework for neural binaural synthesis
Xikun Lu, Fang Liu, Xinni Xie, Ruohan Na, Jinqiu SangNeural binaural synthesis is a promising technique for immersive audio reproduction. However, current methods rely on signal-oriented optimization objectives that treat all time-frequency regions uniformly. This approach conflicts with the selective nature of human auditory perception, as it may overemphasize perceptually masked or sub-threshold components, resulting in audible artifacts in reverberant sound fields. To address this, a perceptual optimization framework inspired by human spatial hearing is introduced. It includes a coherence-gated spatial loss that uses interaural coherence to weight spatial cues according to their reliability. By reducing the contribution of low-coherence regions associated with diffuse reverberation, this mechanism forces the model to prioritize the reconstruction of direct sound and early reflections critical for localization. A multi-scale wavelet loss is also incorporated to measure reconstruction errors across multiple wavelet decomposition levels, addressing the fixed resolution trade-off of the short-time Fourier transform and helping preserve fine-grained temporal transients. The framework was validated on three distinct backbone architectures. Extensive objective evaluations and subjective listening tests under non-head-tracked binaural playback show consistent improvements over state-of-the-art baselines, demonstrating its effectiveness as a model-agnostic training strategy for improving both signal fidelity and spatial accuracy.