Participant-Independent Classification of Autism-Related Visual Attention Patterns from Eye-Tracking Scanpath Images Using a Global–Local Fusion Network
Kun Zhang, Junling Kong, Junhui Zhang, Shuo Zhang, Jingying ChenChildren with autism spectrum disorder (ASD) often exhibit atypical patterns of visual attention allocation and social-cue processing. Eye-tracking scanpath (ETSP) retains information about fixation points, saccade paths and their temporal changes in the form of images, providing an intuitive and computable data representation for analyzing ASD-related visual attention patterns. However, in ASD auxiliary identification studies, the same participant often generates multiple eye-tracking recordings or multiple visual representation samples. If participant independence is not properly considered during model evaluation, the training and test sets may share individualized eye-movement patterns from the same child. In such cases, the model may learn subject-specific characteristics rather than stable and transferable ASD-related visual attention features, leading to an overestimation of its recognition ability on unseen participants. To address this issue, we propose a Global–Local Collaborative Fusion Network (GLCF-Net) under a strict participant-independent splitting protocol. Specifically, the proposed method first maps ETSP images into patch token sequences through a shared Patch Embedding layer. A CNN-based local branch is then used to extract local trajectory morphology, path density, and spatial neighborhood structure, while a ViT-based global branch models cross-region gaze transitions and the overall attention distribution. Finally, a gated adaptive fusion module dynamically integrates local and global information to enhance the representation of stable visual attention features. In the primary repeated stratified five-fold participant-level evaluation, averaging the two out-of-fold probabilities for each participant yielded an Accuracy of 87.0% and a ROC-AUC of 93.7%; the original participant split, retained as a secondary analysis, yielded an Accuracy of 83.52% and a ROC-AUC of 90.27%. Under the reported frozen-backbone configurations, the model also showed a balanced pattern across Accuracy, Recall, and F1-score. These results characterize performance for unseen participants within the same dataset and acquisition conditions.