DOI: 10.1002/ima.70451 ISSN: 0899-9457

Can Annotation‐Free Deep Learning Detect Periapical Bone Rarefactions on Panoramic Radiographs? A Paired Diagnostic Validation Study

José Evando da Silva‐Filho, Caio Marques Silva, Giovanna Tacchi, Igor Rodrigues‐Fontenele, Mateus Bôtto‐Marques, Sâmela Débora Mendes‐Dias, André Wescley Oliveira de Aguiar, Pedro Pedrosa Rebouças Filho, Paulo Leonardo Pontes Marques, Danielle Frota de Albuquerque, Eduardo Diogo Gurgel‐Filho

ABSTRACT

This study aimed to compare an annotation‐free deep learning (DL) model with human evaluators for detecting periapical bone rarefactions (PBRs) on panoramic radiographs (PRs) and to assess the clinical acceptability of AI‐supported dental imaging. A mixed‐methods diagnostic validation was conducted using 238 PRs, equally distributed between images with ( n  = 119) and without ( n  = 119) PBRs. Model performance was compared with that of four calibrated dentists: two oral and maxillofacial radiologists and two endodontists. The reference standard was based on internationally accepted radiographic criteria applied by calibrated researchers under expert supervision, with final classifications verified by the principal investigator. Diagnostic performance was assessed using standard diagnostic metrics, chance‐corrected agreement statistics and receiver operating characteristic analysis, while clinical acceptability was explored through questionnaires and thematically analysed semi‐structured interviews. At the sensitivity‐oriented decision threshold of 0.27, the DL model achieved a sensitivity of 0.924, a specificity of 0.109, an accuracy of 0.517, a precision of 0.509 and an F 1‐score of 0.657. Threshold‐independent binary discrimination was moderate (ROC‐AUC = 0.726). Human evaluators showed more balanced profiles, with accuracy ranging from 0.655 to 0.752 and F 1‐scores from 0.654 to 0.779. Professional perception was generally favourable, although responses varied among evaluators. The annotation‐free DL model showed high sensitivity and moderate discrimination for PBR detection. Although low specificity limits autonomous use, the findings support annotation‐free learning as a potentially scalable strategy, warranting further training with larger and more heterogeneous datasets.