DOI: 10.3390/s26165050 ISSN: 1424-8220

Spatial-Temporal Relation Enhancement for Speech Emotion Recognition from Acoustic Signals Using Fibonacci Encoding with Diverse Feature Fusion

Shicong Huang, Jinghao Zhang, Jingchao Xu, Qi Zhang, Zihan Li, Songyang Wang, Zirui Qiu, Zijia Xiong, Jintian Liang, Changzeng Fu

Speech emotion recognition (SER) infers affective states from speech signals, but positional encoding for acoustic tokens remains underexplored in Transformer-based SER. Existing models often reuse encodings designed for text and do not explicitly account for the different sequential and two-dimensional structures of acoustic representations. We propose Fibonacci Position Embedding (FPE) and Fibonacci Target Shutter (FTS). FTS constructs overlapping candidate-index sets over time–frequency token grids, and FPE samples a Fibonacci index and applies its modulo-wrapped, dimension-dependent phase rotation to query and key vectors. The modules are integrated into STRE-Former, which fuses Wav2Vec, log-mel spectrogram, and MFCC representations through asymmetric cross-representation attention with representation-specific positional encodings. We also introduce an implementation-consistent conditional-entropy formulation that quantifies uncertainty in recovering a token location from its sampled positional representation; this quantity characterizes positional ambiguity rather than downstream modeling capacity. Experiments over 64 positional-encoding combinations on IEMOCAP and MELD identify dataset-dependent highest-observed configurations, reaching 74.21% weighted accuracy on IEMOCAP-4, 74.54% on IEMOCAP-6, and 49.44% on MELD. These empirical observations suggest that the relative behavior of positional-encoding strategies may depend on the acoustic representation and evaluation dataset, rather than supporting a single universally optimal scheme.

More from our Archive