DOI: 10.1002/wics.70077 ISSN: 1939-5108

Machine Learning Methods in Small Area Estimation: A Critical Review of Methods and Applications

Shakeel Ahmed, Muhammad Hamza

ABSTRACT

Small area estimation (SAE) is an important statistical tool applied to produce reliable domain‐level estimates when direct estimation of finite population quantities is not available or unreliable due to limited sample sizes within those domains. While traditional model‐based SAE approaches such as the Fay–Herriot and Battese–Harter–Fuller models possess strong statistical properties, they cannot accommodate ultra‐high‐dimensional auxiliary covariates, high multi‐collinearity, or complex nonlinear relationships. Concurrently, machine learning (ML) algorithmic models have evolved as powerful alternatives which relax rigid parametric assumptions to deliver superior predictive performance. This article provides a comprehensive, critical review of the methods integrating ML algorithms within the SAE landscape across 31 reviewed studies. We introduce a unified three‐tier classification to categorize the literature into: (i) pure predictive ML approaches that completely ignore stochastic random effects, (ii) ML extensions of unit‐level mixed models , and (iii) ML extensions of area‐level mixed models . Utilizing this categorization, we execute a cross‐method synthesis evaluating these architectures across five core pillars: structural fixed‐effects specification, stochastic random‐effects configuration, uncertainty quantification, design‐consistency, and practical interpretability. Our synthesis reveals practical trade‐offs for applied researchers, mapping the data scenarios under which ML methods perform optimally or experience critical inferential challenges, such as random‐effects degeneration under small‐sample domain‐level regimes. Finally, we identify three critical research gaps: the structural lack of survey‐weighted or model‐assisted ML approaches within unit‐level mixed models, the sparsity of area‐level models handling auxiliary measurement errors, and the vulnerability of current nonparametric bootstrap or conformal uncertainty quantification approaches under complex multi‐stage surveys. We conclude by outlining methodological guidelines to guide future research toward developing design‐valid, nonparametric SAE frameworks.