Problem Formulation Outweighs Loss-Function Choice in Deep-Learning Cervical Vertebral Maturation Staging: A Multi-Seed Study on an Imbalanced Public Benchmark
Nazlı Tokatlı(1) Background: Cervical vertebral maturation (CVM) staging on lateral cephalograms informs the timing of orthodontic treatment, and automated deep-learning approaches are widely studied. A recurring but under-examined obstacle is that public CVM datasets are severely imbalanced across the six stages. Using a public dataset and a rigorous eight-seed protocol, we quantify how strongly imbalance constrains performance, whether common imbalance-handling and ordinal-aware loss functions overcome it, and whether a clinically motivated three-group reformulation is more tractable than the native six-stage task. (2) Methods: Using the public Aariz dataset (1000 lateral cephalograms with expert CVM-stage labels; 700/150/150 train/validation/test, one radiograph per patient), we trained an ImageNet-pretrained ResNet-18 under four strategies: unweighted cross-entropy (baseline), class-weighted cross-entropy with balanced sampling, focal loss, and a rank-consistent ordinal (CORAL) head. Each configuration was trained with eight random seeds. Both tasks were evaluated over this multi-seed protocol, and sensitivity analyses examined the ordinal model’s sampling scheme and the focal-loss focusing parameter. Performance was assessed by accuracy, macro-F1, mean absolute error in stages (MAE), and quadratic weighted kappa (QWK), and compared across strategies with the Kruskal–Wallis test; secondary pairwise contrasts were Holm-corrected, and Wilson binomial confidence intervals (accuracy), seed-level confidence intervals (macro-F1), and analytically derived majority and frequency-weighted random classifiers were added as trivial reference baselines. Both the native six-stage task and a three-group scheme (pre-peak = CS1-3, peak = CS4, post-peak = CS5-6) were evaluated. (3) Results: On the six-stage task the extreme imbalance (CS1 n = 18 vs. CS5 n = 311 in training) produced poor minority-stage recognition (macro-F1 ≈ 0.22–0.25), with the rarest stages essentially unclassified even after imbalance handling. Reframing into three clinically meaningful groups roughly doubled macro-F1 (to ≈0.47–0.51) across all strategies. Critically, in the three-group setting no loss function significantly outperformed the others: across eight seeds the Kruskal–Wallis test was non-significant for every metric (macro-F1 p = 0.09, QWK p = 0.05, MAE p = 0.07, accuracy p = 0.41), and the unweighted baseline matched the best specialized loss on macro-F1 (0.507 ± 0.045 vs. ordinal 0.505 ± 0.044). In secondary pairwise comparisons the ordinal loss nominally exceeded the weighted and focal losses (uncorrected p = 0.04 and p = 0.01), but after Holm correction only the comparison with focal loss remained significant (adjusted p = 0.03), and no pairwise difference survived a conservative six-comparison Bonferroni bound; focal loss was the least stable. All strategies clearly exceeded trivial baselines on macro-F1 (majority classifier 0.259; frequency-weighted random ≈0.333), although the majority classifier’s accuracy (0.633) exceeded that of every trained model. (4) Conclusions: On this imbalanced public dataset, the dominant factor associated with CVM classification performance was the problem formulation rather than the loss function: a clinically grounded three-group scheme was far more tractable than six-stage classification, whereas specialized imbalance and ordinal losses did not reliably outperform a simple baseline. These findings, established under a fixed backbone, training budget, and a single public dataset, caution against assuming that loss-level fixes resolve CVM imbalance, and highlight data scale and problem formulation as the primary levers. Multi-dataset external validation is a necessary next step.