Deep Learning for the Assessment of Alveolar Bone Loss on Intraoral Radiographs: A Systematic Review
Nada Tawfig Hashim, Bakri Gobara Gismalla, Muhammed Mustahsen Rahman, Riham Mohammed, Vivek Padmanabhan, Md Sofiqul Islam, Rasha Babiker, Mariam Elsheikh, Ayman Ahmed, Bhavna Jha KukrejaBackground. Radiographic bone loss is a primary determinant of periodontitis stage. Intraoral radiographs (periapical and bitewing) are the standard projections for assessing interproximal bone levels, yet their interpretation is subjective and poorly reproducible, and previous syntheses have pooled intraoral with panoramic imaging. Objectives. To appraise and synthesise studies developing or validating deep learning (DL) models for detecting, quantifying, staging or classifying alveolar bone loss on intraoral radiographs. Methods. Seven databases and six supplementary sources were searched from 1 January 2015 to 12 June 2026 (PROSPERO CRD420261455818, registered retrospectively). Two reviewers screened, extracted and appraised in duplicate using QUADAS-2 with AI-specific signalling questions, CLAIM and APPRAISE-AI; certainty was rated by GRADE. Heterogeneity precluded pooling; synthesis was narrative. Results. Sixteen publications reporting 15 independent datasets (2018–2026) were included (11 periapical, two bitewing, three mixed; 39–21,819 radiographs). The architectures employed comprised classification, segmentation, object-detection, keypoint-localisation and transformer-based models. Accuracy for binary detection ranged from 0.73 to 0.97, Dice coefficients reached ≥0.91, and intraclass correlation with expert measurement was 0.75–0.85. Performance fell for multiclass staging, posterior sites and furcations. Only two studies used an external test set; none was prospective; risk of bias was mostly high or unclear. Conclusions. Performance lies within the range observed for calibrated readers, but the evidence is dominated by small, single-centre, retrospective datasets with annotation-based reference standards and almost no external validation; certainty is very low. Deep learning is best regarded as a clinician-supervised adjunct for screening, triage and quality assurance rather than an autonomous diagnostic device.