Artificial intelligence in radiographic quantification and severity assessment of peri‐implant marginal bone loss: A systematic review
Hooman Khanzadeh, Sanaz Azizigermi, Aida Mokhlesi, Rasoul Gheisari, Gülce Çakmak, Pedro Molinero‐Mourelle, Andrea Roccuzzo, Seyed Ali MosaddadABSTRACT
Purpose
To critically assess artificial intelligence (AI)‐based radiographic models for quantitative measurement, localization/detection, segmentation/keypoints, diagnostic classification, and severity or morphology assessment of peri‐implant marginal bone loss (MBL) and peri‐implantitis‐related bone defects.
Methods
PubMed/MEDLINE, Scopus, Web of Science, Embase, the Cochrane Library, Google Scholar, and reference lists were searched from inception through August 11, 2026. Eligible original studies evaluated AI‐based radiographic assessment of existing dental implants. Quality Assessment of Diagnostic Accuracy Studies‐3 (QUADAS‐3) was applied at the prespecified estimate level for diagnostic/image‐analysis studies and PROBAST for the prediction‐model study.
Results
A total of 1485 records were identified, and 17 studies were included. Fourteen reported localization/detection outcomes, six segmentation/keypoint outcomes, 10 severity/morphology outcomes, 12 diagnostic/classification outcomes, and six direct AI–clinician comparisons; categories overlapped. Implant/peri‐implant tissue detection reached precision of 0.977, recall of 0.992, F1 score of 0.984, and mean intersection over union (IoU) of 0.916. Implant segmentation achieved a Dice of 0.986 and IoU of 0.974, whereas downstream peri‐implantitis classification precision was 0.777. Sensitivity across diagnostic/prediction tasks ranged from approximately 66% to 96%. No study reported the complete prespecified absolute MBL measurement‐agreement outcome set. Six studies used explicitly independent multi‐rater reference standards with consensus and/or reported reliability, and none underwent clearly traceable independent multicenter external validation.
Conclusions
Reported performance is task‐specific and is frequently derived from retrospectively selected, enriched, internally split, or augmented datasets. Current models may support research and carefully supervised radiographic image‐analysis tasks, but none can be recommended for routine clinical use until independent multicenter external validation and prospective studies demonstrate clinically acceptable absolute measurement error and patient‐relevant benefit.