DOI: 10.3390/app16168287 ISSN: 2076-3417

Analyzing Transcript, Emotion-Alignment, and Rhythm Cues for Fake Speech Detection Using Gated Fusion

Alaa Alsaeedi, Amal Almansour, Amani Jamal

Fake-speech detection is commonly studied through acoustic artifacts, speaker-level spoofing cues, or visual inconsistencies, while less attention has been given to meaning-level manipulation, where the spoken content is altered while the speech remains natural and speaker-consistent. This study investigates the relative and complementary contributions of interpretable speech-centered features for detecting meaning-level fake speech. Specifically, it examines transcript-level linguistic and psycholinguistic-style features, text–audio emotion-alignment features, and rhythm-based audio descriptors. A feature-aware gated-fusion framework is used to analyze and combine these three feature groups, with separate branches encoding each feature type and learned branch-level weights adaptively controlling their contributions to binary classification. The framework was evaluated on FakeSpeech+, an audio-only dataset designed for meaning-level manipulation, using a strict leakage-controlled repeated-seed protocol that prevents source-pair, filepath, exact-transcript, and combined group overlap across training, validation, and test partitions. The gated-fusion model achieved 0.842 accuracy, a 0.840 F1-score, and 0.919 AUC. Analysis of the learned fusion weights indicated that transcript-level features contributed most strongly, followed by text–audio emotion-alignment features, while rhythm features received the lowest contribution. These findings provide evidence that meaning-level fake-speech detection can benefit from jointly examining linguistic content, emotional alignment, and rhythmic characteristics, while also highlighting differences in the relative contributions of these feature groups.

More from our Archive