DOI: 10.1111/1755-0998.70190 ISSN: 1755-098X

Accurate Identification of Key Groups of Microeukaryotes Using Multimodal Deep Learning: An Integrated Classification Model Combining Morphological and Molecular Data

Yumeng Song, Lin Zheng, Alan Warren, Mingzhuang Zhu, Bailin Li, Weidong Ji, Xuming Pan

ABSTRACT

Traditional classification of flagellates (i.e., flagellated protists) relies on morphological traits or a single molecular marker, which suffer from subjectivity and limited data sources. This study proposes a multimodal deep learning model, Residual Multi‐Feature Attention‐50 (ResMFA50), that integrates photomicrographs and small subunit ribosomal RNA (SSU rRNA) gene sequences of flagellates. The dual‐branch architecture (DNA sequence and image branches) extracts local and global features, while the Multi‐Feature Attention (MFA) mechanism dynamically fuses heterogeneous data. Experiments were conducted on a dataset comprising 296 SSU rRNA gene sequences and 308 standardized photomicrographs, evaluated using 10‐fold cross‐validation. The results demonstrate that ResMFA50 achieves an accuracy of 92.5% in classifying flagellates at a batch size of eight, which is significantly higher than the accuracies achieved by SVM (82.4%), Random Forest (84.2%), EfficientNet (83.6%), ResNet50 (87.4%), and MMNet (91.3%). Moreover, ablation experiments comparing early, intermediate, and late fusion strategies demonstrate that the proposed late fusion scheme consistently outperforms other fusion timings, achieving improvements of 3.2%–3.8% over early fusion across different batch sizes. This study establishes a methodological foundation for modelling multimodal biological data in complex systems, advancing deep learning applications in integrative taxonomy. This advantage is attributed to the dual‐channel global pooling mechanism (Global Average/Max Pooling fusion), which balances the variance‐bias trade‐off through complementary strategies of spatial statistical smoothing and local salient feature detection, enhancing robustness to data scale expansion.

More from our Archive