Expressive Dizi Synthesis: Unlimited Synthetic Datasets, Real–Synthetic Integration and Technique Labeling
Rongfeng Li, Junchen Liu, Zijin Li, Ya Li, Linfeng Fan, Pei HuangDigital modeling of traditional Chinese musical instruments significantly lags behind that of their Western counterparts, limiting advances in cultural preservation and related research. This paper addresses score-to-audio generation for the Chinese bamboo flute (dizi), aiming to synthesize expressive audio with human-like performance nuances directly from Musical Instrument Digital Interface (MIDI) scores. Two core challenges remain in existing score-to-audio synthesis methods: first, the scarcity of large-scale paired MIDI-audio training data for traditional Chinese instruments; second, the inability of standard MIDI to encode instrument-specific expressive techniques, such as vibrato, pitch bends, and ornamentations. To address these challenges, we propose a scalable workflow that generates large-scale synthetic MIDI-audio pairs through rule-based score randomization, automated digital audio workstation (DAW) rendering with commercial sample libraries, and standardized feature extraction. Performance technique information is directly encoded into the MIDI stream using out-of-range MIDI note numbers, achieving more effective conditioning than external control methods. Our system is built upon the Musical Instrument Digital Interface–Differentiable Digital Signal Processing (MIDI-DDSP) framework. Large-scale synthetic data is used for pre-training to establish timbral consistency, while real recordings from the University of Rochester Multi-Modal Music Performance (URMP) flute corpus are employed for fine adjustment to achieve expressive dynamic variations. Synthetic data alone yields stable but less expressive outputs, whereas real data alone risks overfitting. The combined strategy achieves a realistic timbre with controllable dynamics and performance techniques.