DOI: 10.1002/sam.70105 ISSN: 1932-1864

Testing for Conformity With Parametric Distributions Using Level Sets: From Univariate to Bivariate Applications

Rasha Alsaadawi, Nitai D. Mukhopadhyay

ABSTRACT

Conformity of observed data with a postulated probability distribution is a fundamental problem of statistical modeling. Current tools for checking distributional conformity are often limited to univariate settings, with extensions to multivariate distributions facing challenges in interpretability and power. We propose a novel topological data analysis (TDA) framework that uses level sets to test distributional conformity, extending from univariate to bivariate applications. Level sets provide a natural geometric representation of probability distributions, capturing essential shape characteristics across dimensions. For univariate data, we construct level sets and compare them using the Dice Similarity Coefficient (DSC) to build a measure of distributional conformity; we then extend this methodology to bivariate settings through kernel density estimation and an annular decomposition of probability‐based level sets, with inference via an adaptive permutation test. Across simulation studies, the method maintains nominal Type I error control and achieves power competitive with established tests, including the Shapiro–Wilk test in the univariate case and the energy‐based E‐statistic and Henze–Zirkler test in the bivariate case, though no single method dominates uniformly across all alternatives. Rather than offering uniformly greater power, the framework is particularly well suited to settings where the goal is to characterize where and how distributions differ: the level set decomposition identifies which density regions drive a detected difference, and the average DSC provides a bounded, interpretable effect size on the unit interval. We illustrate these capabilities through applications to NHANES data, comparing cholesterol distributions and the joint distribution of systolic blood pressure and body mass index across age groups. This geometric interpretability, combined with natural extensibility to higher dimensions, positions our method as a valuable complement to existing approaches in the modern statistical toolkit.

More from our Archive