DOI: 10.26599/bsa.2026.905026 ISSN: 2096-5958

CDESD: The Chinese Disyllabic Emotional Speech Database

Licheng Mo, Sijin Li, Jiahong Tang, Yaohua Dong, Dandan Zhang

Background:

Existing emotional speech corpora do not always meet the need for brief, systematically characterized Mandarin stimuli in cognitive and affective neuroscience research. This study developed the Chinese Disyllabic Emotional Speech Database (CDESD) to address these methodological gaps.

Methods:

Three native female Mandarin speakers produced 85 disyllabic stimuli under three emotional conditions: neutral, angry, and happy. Recordings were standardized to 1-s files with a consistent nominal onset structure. Ten participants evaluated emotion recognition accuracy and expression intensity. Core acoustic features—including voiced duration, fundamental frequency (F0), sound intensity, and harmonic-to-noise ratio—were extracted using Python-based analysis tools. An independent sample of 37 participants provided valence, arousal, and dominance ratings.

Results:

The final database comprised 765 recordings, with 255 per emotion category. Mean emotion recognition accuracy ranged from 89% to 91%. Emotion categories differed significantly in valence, arousal, and dominance. Happy stimuli received the highest valence and dominance ratings, whereas angry stimuli received the highest arousal ratings. However, dominance ratings for happy and neutral stimuli showed negligible inter-rater reliability and should be interpreted with caution.

Conclusions:

The CDESD provides a systematically characterized stimulus resource for emotional speech processing research, particularly for studies requiring brief Mandarin disyllabic materials. Its fixed file duration and standardized nominal onset structure facilitate consistent stimulus presentation in event-related potential paradigms, while accompanying item-level acoustic annotations allow researchers to select matched subsets or statistically control for relevant acoustic differences.