DOI: 10.1093/jrsssa/qnag095 ISSN: 0964-1998

Fake news detection via wisdom of synthetic and representative crowds

François t’Serstevens, Roberto Cerina, Giulia Piccillo

Abstract

Social-media companies have struggled to define fake news in a democratically legitimate way. Reliance on expert judgement has faced criticism due to trust deficits and political polarization. Crowd-based approaches offer a cost-effective, transparent, and inclusive alternative. This article presents a novel end-to-end method to detect fake news on X using the wisdom of synthetic, representative crowds. Drawing from multilevel regression and poststratification literature, we train a Hierarchical Bayesian model to predict tweet veracity from the perspective of diverse personae within the target population. We then weight these predictions using a representative stratification frame, ensuring decisions about ‘fake’ tweets reflect the broader polity. Based on the aggregated scores, we analyse a corpus of tweets and conduct a second multilevel regression and poststratification to estimate the characteristics of fake news sharers. We find small but statistically meaningful heterogeneity in US fake news sharing on X: (i) sharing is rare, with a low average probability; (ii) strong evidence that Democrats share less fake news than Republicans; and (iii) when Republican definitions are used, Republicans show a decreased propensity to share fake news.

More from our Archive