DOI: 10.3390/electronics15194486 ISSN: 2079-9292

Cleaning Contaminated Reference Sets for Few-Shot Visual Anomaly Detection

Sergio Villanueva López, Emilio Soria-Olivas, Manuel Sánchez-Montañés

Few-shot anomaly detectors for industrial inspection typically use between five and 20 reference images to learn what a normal part looks like. These images are assumed to be defect-free, which is not 100% guaranteed on an actual production line. This makes the reference set a potential point of failure: if an image of a defective part ends up among the references, the detector will silently accept similar defects. First, we quantify this degradation using 35 categories from the public datasets MVTec AD, VisA, BTAD, and MVTec LOCO. We show that replacing just three of 10 reference images with defective parts results in a loss of 1.8 percentage points of AUROC on average, and up to 3.0 depending on the dataset. This represents a substantial loss for an industrial anomaly detection system. Second, we present an unsupervised audit algorithm that identifies which reference images are likely defective, and requires no training. The audit is a one-time preprocessing step that evaluates the internal consistency of the reference set before it is used for inspection. Each image is compared to the rest of the references, scored based on its most atypical region, and flagged when that score stands out. A technician can then remove, review, or recapture the flagged images before the system is deployed on the production line. We show that recapture recovers more than half of the accuracy lost due to contamination, and if there are no truly defective images, it does not degrade the system’s performance. Our method is intended for memory-bank detectors built from a handful of references, where an isolated part with a structural or textural defect can slip into the set. It runs in seconds on CPU, makes its decisions at the image level independently of the anomaly detection algorithm used, and can be easily incorporated into existing systems. The code and results are publicly available.