DOI: 10.14778/3819518.3819564 ISSN: 2150-8097

Finding Non-Redundant Simpson's Paradox in Multidimensional Data

Yi Yang, Jian Pei, Jun Yang, Jichun Xie

Simpson's paradox has broad impact across many scientific domains. Existing detection methods overlook a key issue: many detected paradoxes may be redundant , arising from equivalent data subsets, identical subpopulation partitions, or correlated outcome variables, thereby obscure insights and increase computational cost. In this paper, we present a framework for finding non-redundant Simpson's paradoxes by formalizing three sources of redundancy—sibling child, separator, and statistic equivalence—and showing that pair-wise redundancy forms an equivalence relation. We further propose a concise representation that groups redundant paradoxes and develop efficient algorithms combining depth-first population materialization with redundancy-aware discovery. Experiments on real and synthetic datasets show that redundancy is prevalent (over 40% in some cases), while our methods scale to millions of records, achieve up to 6.72× speedup over brute-force approaches and identify robust paradoxes, enabling efficient discovery, compact summarization, and clear interpretation in multidimensional data.

More from our Archive