Information-Preserving and Model-Aware Voronoi Dequantization of Weighted Discrete Laws
Maha Moussa, Khater A. E. Gad, H. M. Hamouda, Mohamed F. Abouelenein, Ahmed Sedky EldeebContinuous dequantization embeds discrete data into a continuous space, but relaxation can alter the statistical information carried by the original categories. We study a complementary regime in which a specified weighted discrete law is the target and the dequantizer is required to be lossless under a prescribed quantizer. Any partition-respecting, parameter-independent kernel produces a continuous experiment Blackwell equivalent to the original discrete one, so likelihood ratios, maximum-likelihood estimators, score and Fisher information, Bayes posteriors, and optimal risks are unchanged before downstream approximation. We then characterize what happens after a finite-capacity continuous model is fitted. A Kullback–Leibler (KL) chain rule separates continuous approximation error into the categorical error of the requantized model plus a within-cell shape term; the information-preservation guarantee therefore does not extend automatically to an arbitrary fitted downstream model. We also provide a plug-in analysis for empirically estimated weights, showing that dequantization preserves rather than removes the discrete estimation error. Within each cell, maximizing entropy minus a displacement penalty yields a truncated Gibbs kernel. The apparent concentration and geometric-scale parameters reduce exactly to one effective concentration, α=λ/sc2. We further derive exact transport costs, a sharp W1=Θ(h2) reconstruction rate for the special reconstruction-from-exact-bin-masses setting, leakage-specific and geometry-perturbation bounds, and a bounded-domain multidimensional formulation. A finite-candidate Monte Carlo selector satisfies a nonasymptotic O(m−1/2) oracle inequality. Experiments verify the identities and show that exact dequantization can materially reduce requantized error when a smooth continuous downstream model is imposed; this is a conditional representational advantage, not a claim that dequantization dominates an unrestricted categorical model.