DOI: 10.1145/3839564 ISSN: 2770-6699

Measuring What Matters: Consistency and Compactness in Evaluation of Counterfactual Explanations

Amir Reza Mohammadi, Andreas Peintner, Michael Mueller, Eva Zangerle

Explainability in recommender systems (RS) remains a pivotal challenge. Counterfactual explanations have emerged as a particularly actionable paradigm, offering intuitive “what-if” reasoning. However, their evaluation lacks principled standards. Current metrics primarily assess whether explanations change the top-ranked recommendation, overlooking two fundamental aspects of explanation quality. First, evaluation results are frequently inconsistent , as metrics are tightly coupled to the underlying recommender’s performance. Second, explanations are rarely assessed for compactness —whether changes are sufficiently small to remain interpretable. Oversized counterfactuals may technically succeed but fail to provide practical insight.

In this work, we advocate for a holistic evaluation perspective centered on consistency and compactness . We systematically analyze how extending evaluation beyond top-1 to top-k recommendations improves metric stability and reduces dependence on recommender fluctuations. In parallel, we introduce compactness-aware evaluation criteria that quantify the minimality of counterfactual modifications. Through extensive experiments across multiple datasets and models, we demonstrate that jointly considering these dimensions yields more reliable assessments. Our findings expose key factors driving metric instability, highlight the trade-off between effectiveness and explanation size, and provide practical guidelines toward standardized, faithful, and compact evaluation of counterfactual explanations in recommender systems.

More from our Archive