DOI: 10.3390/electronics15184293 ISSN: 2079-9292

Relation Prototype Re-Scoring for CLIP-Based Logical Anomaly Detection and Localization

Hanhoon Park

CLIP-based anomaly detectors have markedly advanced training-free and zero-shot industrial anomaly detection and localization, yet their predictions remain dominated by patch-wise vision–language similarity or anomaly-aware feature scoring. This formulation is intrinsically limited for logical anomalies, in which every visible part can appear locally normal while its count, position, arrangement, or co-occurrence violates a normal configuration. We introduce a training-free relation prototype re-scoring module that reuses the semantic–spatial relations already encoded by the visual transformer. Because patch tokens contain positional embeddings, the self-attention graph is position-aware as well as content-dependent; we use it as a message-passing operator over detector-specific patch features. Normal images define category-wise, spatially indexed relation prototypes, and each test patch is scored by the Euclidean deviation of its attention-aggregated feature from the corresponding normal prototype. The resulting map supports both dense localization and map-derived image detection. Although we also report the feature-relation map alone to analyze its intrinsic behavior, the final method fuses this map with the original anomaly map so that relation-sensitive evidence is added without discarding the baseline detector’s local appearance cues. We develop the method on AnomalyCLIP and verify its generality on WinCLIP and AA-CLIP. On the logical split of MVTec LOCO AD, the proposed fusion improves AnomalyCLIP from 53.9 to 74.4 pixel AUROC and from 28.7 to 52.7 pixel AUPRO, while image AUROC rises from 50.5 to 71.7. Similar improvements are observed for WinCLIP and AA-CLIP. On CAD-SD, the proposed fusion reaches 92.0 and 93.8 image AUROC for AnomalyCLIP and WinCLIP, respectively, on co-occurrence anomalies. Experiments on the structural split of MVTec LOCO AD and the MVTec AD benchmark show why the baseline map must be retained: fusion preserves substantially more local-defect evidence than relation-only scoring, although its benefit remains detector- and category-dependent. These results identify semantic–spatial relation deviation as a missing cue in CLIP-based logical anomaly detection and localization, without requiring explicit component models, symbolic rules, additional training, or modification of the baseline architecture.