DOI: 10.1098/rsos.251703 ISSN: 2054-5703

Making hidden biases visible in population location data from mobile phones

Carmen Cabrera, Francisco Rowe

Abstract

Traditional sources of population data, such as censuses and surveys, are costly, infrequent and often unavailable in crisis-affected regions. Mobile phone data (MPD) offer near-real-time, high-resolution insights into population distribution, but their utility is undermined by unequal access to digital technologies, creating biases that threaten representativeness. Despite recognition of these issues, no standard framework exists to address such biases. We develop and implement a systematic, replicable framework to quantify and explain population coverage bias in aggregated MPD without requiring individual-level attributes. The approach combines an indicator of population coverage bias with explainable machine learning to identify contextual drivers of spatial variation in bias. Using four datasets for the UK benchmarked against the 2021 census, we show that MPD achieve higher population coverage than national surveys, but biases persist across sources and subnational areas. Population coverage bias is associated with demographic, socio-economic and geographic features, often in complex nonlinear ways. Contrary to common assumptions, multi-application datasets do not necessarily reduce bias compared to single-app sources. By providing a transferable framework for measuring and explaining population coverage bias, our study advances the methodological infrastructure required to evaluate, compare and responsibly integrate MPD into population research, official statistics and evidence-based policy.

More from our Archive