MoR–Swin: Efficient Vision Transformer Using Mixture of Recursions
Yongbao Ai, Tianxiang Gao, Zhipeng Lin, Longqi Yang, Qingyu ChangVision Transformers, especially Swin Transformer, have become default backbones for various vision tasks but suffer from high memory consumption and training costs. This letter proposes MoR–Swin, a novel architecture that integrates Mixture of Recursions (MoR) into Swin Transformer. An adaptive token-level recursion mechanism dynamically allocates computational depth based on semantic complexity. A recursive window attention module and a lightweight router with load balancing loss are introduced. Extensive experiments on ImageNet classification, COCO detection, and ADE20K segmentation show that MoR–Swin reduces parameters by about 50% and accelerates inference up to twofold at a modest accuracy cost (within about 0.5 points of Swin-B on ImageNet-1K). It provides a new technical pathway for optimizing Vision Transformer models, significantly enhancing their applicability in resource-constrained environments.