RWEB-CWBI: A Reliability-Weighted, Empirical-Bayes-Shrunk Composite Estimator for Multidimensional Child Well-Being, with Cross-Survey Validation in Egypt
Mariam Magdy Kamal, Mohamed Naguib, Noura Anwar Abdel-FatahBackground: Composite well-being indicators are typically built by aggregating dimension scores with subjectively elicited or transferred weights, computed on raw sample means without formal reliability checks. Methods: We propose the Reliability-Weighted, Empirical-Bayes-Shrunk Composite (RWEB-CWBI), which shrinks each dimension’s group-level mean via the dimension-specific DerSimonian–Laird random-effects models, derives aggregation weights from a group-level principal component analysis of the shrunk dimension matrix, and rescales each loading by its empirical area-level reliability ratio (RR), automatically reducing the influence of dimensions with lower estimated reliability, and thus providing an alternative to automatic exclusion. RWEB-CWBI is applied within a broader Child Well-Being Index (CWBI) framework, building on participatory CWBI1-CWBI5 specifications estimated on a 10-governorate Takaful-beneficiary sample (n = 8716 households; the 2016 UNICEF sample), and evaluated alongside data-driven weighting, empirical Bayes shrinkage, and machine-learning benchmarking on an independent, nationally representative sample (EFHS-2021; n = 49,936 children, 26 governorates). Results: Participatory CWBI4 and CWBI5 showed the strongest alignment with the available subjective well-being measures. RWEB-CWBI down-weighted the least reliable EFHS dimension (RR = 0.75) automatically, while preserving close agreement with the simpler indices’ governorate rankings (ρ = 0.988–0.989). Gradient boosting improved cross-validated R2 over OLS by 2.5 points (0.798 vs. 0.773), with urban residence dominating permutation importance. The EFHS-UNICEF ranking correlation stayed weak and non-significant under every specification, including RWEB-CWBI (ρ = 0.14–0.19, n = 10, p ≥ 0.60). Conclusion: RWEB-CWBI is presented as a statistically motivated methodological proposal for handling heterogeneous dimension-level precision and availability. The persistence of the weak EFHS-UNICEF correlation across conventional, data-driven, and reliability-weighted specifications indicates that the finding is robust to the choice of weighting method, although contributions from population differences, measurement differences, and survey-design differences cannot be distinguished conclusively.