DOI: 10.3390/e28101073 ISSN: 1099-4300

Sample-Conditional Mixtures of Entropy Estimators for Short Sequences Under a Uniform Marginal Null

Guillermo Sosa-Gómez

Entropy estimation from short samples recurs in symmetric cryptography, where the reference distribution is uniform by design and no single estimator in the considered classical comparison set minimizes mean squared error (MSE) across the full range of ratios n/k. We introduce the adaptive sample-conditional entropy diagnostic (ASED), in which a compact network trained offline maps a frequency-of-frequencies descriptor of the sample to convex mixture weights over six classical estimators, at a cost of O(n+k). We prove an oracle inequality bounding the excess risk of such a mixture by the L1 error of its weights and fixed-alphabet consistency for every simplex-valued weighting rule, independently of the distribution used to train the weights. Under the uniform-null protocol, ASED attains an integrated MSE of 1.7×10−3 for bytes, compared with 3.8×10−2 for James–Stein shrinkage, a reduction that depends materially on the aggregation scheme: 95.4% for summed or averaged MSE across sample sizes versus 24% for the mean of per-size ratios. Both figures are reported together throughout, and neither is presented as the headline; the entire advantage is confined to the undersampled regime n<k, the two estimators being indistinguishable for n≥k. A locked train/validate/test evaluation with an independently written implementation reproduces the reduction at 95.1%. Because a constant output achieves zero error by construction here, we add further checks: in a post hoc analysis, ASED is non-inferior to SHR in detection power at a 0.02 margin, both pointwise and simultaneously, whereas a constant control has none, and the advantage persists on held-out configurations. The 0.02 margin was selected after inspecting the power estimates, and at n=8 and n=16, it is smaller than the variation induced by tie handling at the empirical critical value, so the non-inferiority conclusion is confirmatory only for n≥32. At n≤16, the two statistics have identical null distributions up to monotone relabeling, and no test of size 0.05 exists, so no power difference is identifiable in either direction, and the comparison is withdrawn there. For strongly non-uniform sources, the ordering reverses, and NSB dominates, delimiting ASED as a uniformity diagnostic rather than a general-purpose entropy estimator. The design choices behind ASED, the feature map, architecture, and loss weighting were made while observing results on this same evaluation protocol, so no independent model-selection split separates development from the assessment reported here. That pipeline is accordingly designated exploratory, and the locked evaluation is confirmatory. The contribution is stated as sample-conditional convex aggregation of classical estimators rather than as the particular network realizing it: a five-feature variant and a gradient-boosted surrogate perform comparably, with no paired interval separating the three. Claims are restricted to short i.i.d. samples and marginal Shannon-entropy diagnostics under a uniform null; min-entropy, unpredictability, dependence, and randomness certification are out of scope.