Hybrid Landmark-Guided and EfficientNet Feature Fusion for Down Syndrome Facial Screening
Meshal Alfuraydi, Hassan MathkourDown syndrome is associated with characteristic craniofacial features that have motivated the development of computer-vision-based facial-image-based screening systems. Recent studies have increasingly relied on deep computer vision and deep-learning approaches, but many provide limited interpretability, while earlier landmark-based methods offered transparent geometric and texture-based measurements. This creates a gap between interpretable handcrafted features and high-performing deep representations. To address this gap, this study proposes a hybrid interpretable–deep framework that combines landmark-derived geometry features, landmark-guided local binary pattern (LBP) texture descriptors, and frozen EfficientNetB0 convolutional neural network (CNN) deep features. The primary contribution of this study is the systematic integration and comprehensive evaluation of complementary interpretable and deep-feature representations within a unified facial-image screening framework. Feature fusion is followed by random forest feature ranking and SVM-RBF classification. Experiments were conducted on 2979 successfully processed facial images from an original dataset of 2999 images. Geometry-only, texture-enhanced, deep-feature, and hybrid fusion models were evaluated using repeated stratified train–test splits. The final RF Top-800 fusion model achieved strong facial-image classification screening performance, with F1 = 0.9045 +/− 0.0134 and AUC = 0.9675 +/− 0.0073 across repeated stratified train–test splits. Ablation analysis showed that removing geometry features, removing LBP features, or using only EfficientNetB0 reduced performance, supporting complementary contributions from interpretable geometry and texture feature components and deep-feature representations. Statistical comparisons and duplicate-sensitivity analyses further supported the robustness of the results. The findings demonstrate that landmark-derived geometry and texture descriptors remain valuable when integrated with modern deep representations, providing feature-level interpretability while improving screening performance through a hybrid framework that combines interpretable handcrafted features with high-performing deep representations.