DOI: 10.1145/3839564 ISSN: 2770-6699
Measuring What Matters: Consistency and Compactness in Evaluation of Counterfactual Explanations
Amir Reza Mohammadi, Andreas Peintner, Michael Mueller, Eva Zangerle
Explainability in recommender systems (RS) remains a pivotal challenge. Counterfactual explanations have emerged as a particularly actionable paradigm, offering intuitive “what-if” reasoning. However, their evaluation lacks principled standards. Current metrics primarily assess whether explanations change the top-ranked recommendation, overlooking two fundamental aspects of explanation quality. First, evaluation results are frequently
inconsistent
, as metrics are tightly coupled to the underlying recommender’s performance. Second, explanations are rarely assessed for
compactness
—whether changes are sufficiently small to remain interpretable. Oversized counterfactuals may technically succeed but fail to provide practical insight.
In this work, we advocate for a holistic evaluation perspective centered on
consistency
and
compactness
. We systematically analyze how extending evaluation beyond top-1 to top-k recommendations improves metric stability and reduces dependence on recommender fluctuations. In parallel, we introduce compactness-aware evaluation criteria that quantify the minimality of counterfactual modifications. Through extensive experiments across multiple datasets and models, we demonstrate that jointly considering these dimensions yields more reliable assessments. Our findings expose key factors driving metric instability, highlight the trade-off between effectiveness and explanation size, and provide practical guidelines toward standardized, faithful, and compact evaluation of counterfactual explanations in recommender systems.