DOI: 10.3390/s26165101 ISSN: 1424-8220

Design and Efficacy of a Speaker Verification Method Combining CNN and Transformer for Secure Access Control

Xiao Li, Xiao Hu, Kun Niu, Ling Tuo

In the rapidly evolving landscape of computer and mobile applications, the demand for secure access control has become increasingly pivotal. This paper introduces a speaker verification method aimed at remotely verifying an individual’s claimed identity, so as to achieve access authorization. The primary objective is to develop a deep learning network capable of eliminating redundant and irrelevant information while learning robust deep speaker embedding descriptors that capture speaker-specific idiosyncrasies. This paper presents an AI framework combining convolutional neural network and transformer architectures, enhanced by (1) a stereoscopic attention mechanism that computes fine-grained attention weights across frequency, time, and channel dimensions with fewer parameters than existing CBAM or SE modules, and (2) a multi-feature aggregation mechanism that fuses supervised and unsupervised features at the utterance level to maximize complementary speaker information. These innovations introduce a new perspective for integrating acoustic and articulatory features, effectively addressing the challenges related to short-segment speech and cross-domain scenarios. Experimental results demonstrate that the AI framework achieves competitive performance under domain mismatch conditions. The methodology acts as a catalyst for advancing application authorization, highlighting the transformative potential of AI-driven innovations in the software engineering field.

More from our Archive