Comparative Evaluation of CNN, Vision Transformer, and Swin Transformer Feature Representations with CatBoost for Bethesda Classification of Multi-Cell Cervical Cytology Image Patches
Miguel Angel Valles-Coral, Ciro Rodriguez, Carlos Navarro, Ulises Román-Concha, Diego RodriguezArtificial intelligence in cervical cytology requires controlled comparisons to distinguish the contribution of visual representations within the classification pipeline. This study evaluated six deep feature representation architectures for Bethesda classification of nucleus-centered multi-cell cervical cytology patches. The CRIC Cervix Collection comprised 400 whole-slide images (WSIs) and 11,530 annotated cellular records across six Bethesda categories. DenseNet121, DINOv2 ViT-B/14, Swin Transformer Tiny, VMamba-T, GFNet-S, and MLP-Mixer-B/16 were evaluated under frozen and partial fine-tuning conditions using CatBoost as a common downstream classifier and WSI-level partitioning to reduce data leakage. In the ten-fold post-selection Dev80 evaluation, DenseNet121 achieved the highest mean Macro-F1 under partial fine-tuning (0.5038 ± 0.0318) and frozen extraction (0.4858 ± 0.0305). However, on the predefined locked Test10 evaluation after Final90 training, partially fine-tuned DINOv2 achieved the strongest overall performance, with a Macro-F1 of 0.4965, Balanced Accuracy of 0.5563, Macro ROC-AUC of 0.9155, and LogLoss of 0.8579. Exploratory Friedman analyses indicated differences among representations, although DenseNet121 and DINOv2 were not statistically distinguished after Holm correction. These findings show that internal stability and locked-test performance may yield different model rankings and support leakage-aware evaluation of heterogeneous visual representations in cervical cytology.