DAPR: Dynamic Distribution-Aware and Adaptive Pseudo-Label Refinement for Long-Tailed Semi-Supervised Oral Disease Classification
Xuesheng Bian, Zeyu Xie, Yuhan Sun, Jinxin Zhu, Wei Li, Hong JiaOral disease image classification can support computer-assisted assessment of intraoral images. However, obtaining large-scale annotated medical data is expensive, while real-world oral disease datasets often exhibit severe long-tailed distributions, where minority disease categories contain only limited samples. Existing semi-supervised learning methods commonly rely on fixed-threshold pseudo labels and may produce prediction distributions dominated by majority classes, resulting in class bias and pseudo-label noise accumulation under long-tailed settings. To address these issues, we propose DAPR, a Distribution-Aware and Adaptive Pseudo-Label Refinement framework for long-tailed semi-supervised oral disease classification. DAPR employs Dynamic Category Distribution Modeling (DCDM) to track the evolving prediction distribution of unlabeled samples and generate distribution-aware soft pseudo labels. A class-adaptive dynamic thresholding (CADT) mechanism was further introduced to improve minority-class sample utilization. In addition, Relation-aware Representation Learning (RRL) aligns semantic and feature relationships to enhance feature discrimination. Experiments using stratified five-fold cross-validation on a long-tailed oral disease dataset show that DAPR achieves the highest average Accuracy and Macro-F1 among the compared methods under the evaluated setting. DAPR achieves the highest average Accuracy and Macro-F1 among the compared methods and obtains strong aggregate tail-class performance, particularly for Tooth Discoloration and Ulcers. These results indicate that DAPR improves aggregate class-balanced learning under the evaluated dataset, annotation ratio, and backbone configuration.