DOI: 10.3390/diagnostics16162569 ISSN: 2075-4418

Algorithmovigilance in AI-Based Oral-Health Surveillance: Temporal Drift, Cross-Survey Differences, and Predictor-Level Structural Stability in Population-Level Severe Tooth Loss Prediction

Quang Tuan Lam, Fang-Yu Fan, Yung-Li Wang, Sheng-Wei Feng, Thi Thuy Tien Vo, Minh Huu Nhat Le, Giang Vu, Nguyen Quoc Khanh Le, I-Ta Lee

Background/Objectives: Artificial intelligence models may retain acceptable discrimination while calibration and predictor–risk relationships change across populations and data-collection systems. We evaluated a structured algorithmovigilance framework for detecting temporal drift, cross-survey differences, and predictor-level structural instability in severe tooth loss prediction. Methods: Across five BRFSS cycles from 2016 to 2024 (total N = 2,176,039), a survey-weighted main-effects Explainable Boosting Machine trained in 2016 was evaluated chronologically through 2024. NHANES 2015–2018 served as an external reference for an exploratory cross-survey temporal comparison. Results: AUC declined modestly from 0.8638 in 2016 to 0.8495 in 2024. The frozen unrecalibrated model yielded AUC = 0.8681 and Brier Score = 0.1131 in pooled NHANES. The survey-system-by-period interaction was positive (beta = 0.0412; HC1 p = 0.014), although a stratified survey-weighted bootstrap sensitivity produced a wider interval including zero (95% CI, −0.0085 to 0.0895; p = 0.093). Smoking (p < 0.001) and income (p = 0.030) showed nominal predictor-level structural shifts, whereas BMI did not. An interaction-enabled EBM improved AUC by 0.0010 in 2016 and 0.0016 in 2024. Conclusions: Aggregate discrimination alone did not capture calibration and predictor-level changes. The framework provides retrospective monitoring signals for governed review, not evidence of causal survey-mode effects, formal algorithmic fairness, or clinical deployment readiness.

More from our Archive