DOI: 10.3390/diagnostics16193132 ISSN: 2075-4418

A Clinician-Configurable FAHP-Based Ensemble Framework for Multi-Label Chest X-Ray Classification

Hui-Chu Chiu, Yu-Hsiang Tsai, Cheng-Hsuan Juan, Chia-Ching Chang, Ya-Hui Li, Deng-Yiu Chiu, Chun-Jung Juan

Background/Objectives: Deep learning-based chest X-ray (CXR) analysis has shown strong diagnostic performance; however, conventional ensemble methods typically rely on equal-weight aggregation and do not account for disease-specific model behavior or clinician-defined priorities. Methods: We developed a clinician-configurable ensemble framework based on the fuzzy analytic hierarchy process (FAHP) for multi-label CXR classification. The proposed method employs a two-stage design consisting of disease-specific model weighting using multi-criteria decision analysis and utility-guided seed aggregation based on ΔNet Benefit (ΔNB). Five base architectures, including Vision Transformer (ViT), SwinV2 Transformer, DenseNet121, DenseNet121 + ViT and DenseNet121 + SwinV2, were trained on the National Institutes of Health (NIH) ChestX-ray14 dataset using 10 random seeds. Performance was evaluated on a held-out test set using area under the receiver operating characteristic curve (AUROC), area under the precision–recall curve (AUPRC), balanced accuracy, and F1 score and compared with single-model baselines, Hard Voting, and Soft Voting. Results: In the seed-level analyses, Soft Voting achieved the highest overall discrimination performance (AUROC, 0.851 ± 0.001; AUPRC, 0.293 ± 0.003), whereas seed-level FAHP showed comparable performance (AUROC, 0.850 ± 0.001; AUPRC, 0.287 ± 0.004). In the final-level analyses, the sensitivity-priority FAHP–ΔNB configuration achieved a macro-averaged AUROC of 0.855 and an AUPRC of 0.299. Hard Voting demonstrated inferior performance, particularly for AUPRC. Per-disease analysis revealed heterogeneous results, with FAHP showing favorable performance in selected conditions, including Effusion, Emphysema, Mass, Nodule, and Pleural Thickening, while Soft Voting remained superior in others. Statistical analysis confirmed that differences between FAHP and Soft Voting were generally modest and disease-dependent. Conclusions: The proposed FAHP-based ensemble maintains performance comparable to strong equal-weight baselines while enabling clinician-configurable, disease-specific model weighting and utility-guided aggregation. This framework provides an aggregation-level transparent approach for incorporating configurable diagnostic priorities into multi-label CXR classification.