DOI: 10.1111/1556-4029.70432 ISSN: 0022-1198

EMBNet: Multi‐scale feature learning with efficient channel attention for deepfake speech detection

Haitao Yang, Fen Li, Xin Cai, Juanjuan Huang, Yifan Tang, Yingzhuo Xiong, Huapeng Wang

Abstract

With the rapid advancement of generative artificial intelligence, deepfake speech has emerged as a significant threat to digital audio authenticity, posing challenges for forensic analysis and legal applications. In this study, we propose EMBNet, a task‐oriented deepfake speech detection framework that integrates efficient channel attention (ECA) with a multi‐scale bottleneck to enhance the representation of subtle acoustic anomalies. The ECA module adaptively emphasizes critical feature channels, while the multi‐scale bottleneck captures local and hierarchical spoofing traces across multiple time–frequency resolutions. The proposed framework effectively balances fine‐grained local detail modeling with global hierarchical representation, improving the detection of weakly manifested spoofing patterns. Extensive experiments on the ASVspoof 2019 Logical Access dataset demonstrate that EMBNet significantly outperforms existing baseline and state‐of‐the‐art methods, achieving an EER of 2.67%, an AUC of 97.32%, and an F1‐score of 97.28%. Ablation studies further confirm the complementary contributions of the ECA and multi‐scale modules to overall performance. The proposed approach demonstrates promising performance for forensic audio analysis and may support forensic audio authenticity assessment by improving the discrimination between genuine and manipulated speech.

More from our Archive