DOI: 10.3390/make8100304 ISSN: 2504-4990

Contrastive Gradient-Flow Interpretation (CGFI): A Composable Operator for Class-Discriminative Explanation of Convolutional Neural Networks

Fatemeh Barati, Mohammad Soltanian, Keivan Borna, Luca Longo

Growing deployment of Convolutional Neural Networks in safety-critical domains has intensified the need for transparent, discriminative explanations, yet gradient-based methods such as Grad-CAM compute explanations for a single class score, producing saliency maps that activate regions shared with competing alternatives. This limitation is particularly consequential in multi-class and fine-grained recognition settings, where non-discriminative explanations reduce the practical utility of the method and impede reliable model auditing. This research study contributes to the body of knowledge by proposing Contrastive Gradient-Flow Interpretation (CGFI), an operator on gradient-weighted attribution maps that applies the contrastive principle that discrimination is sharpened by modelling what a class is not, to generate class-discriminative explanatory maps. CGFI explicitly models the top-K competing classes, computes their aggregated gradient attribution, and subtracts their influence from the target class attribution at the feature-map level, emphasising regions uniquely characteristic of the predicted class. It was evaluated across four backbones (Grad-CAM, Grad-CAM++, XGrad-CAM, Score-CAM), with and without contrast, on ResNet-50, VGG-16 and VGG-19. CGFI reduced the target-rival rank correlation across all 1162 image-backbone combinations evaluated, and improved the target probability share in 11 of 12, while preserving attribution fidelity. These findings suggest that contrastive attribution is a principled mechanism for improving explanation specificity.