DOI: 10.1177/01979183261473318 ISSN: 0197-9183

Analyzing Online Migration Forums: An Introduction to Natural Language Processing for International Migration Research

Nari Yoo, Donghun Kim, Sou Hyun Jang, Eric Fong

In contemporary migration studies, online forums are common spaces where people exchange information and seek advice about international migration. This methods note introduces the Reddit community r/IWantOut as a data source for studying migration aspirations and presents a validated natural language processing (NLP) pipeline for analyzing it. The forum's community rules require posters to encode age, gender, origin, and intended destinations in a fixed title format, so users label their own migration aspirations in a semi-machine-readable form when they post. Using 156,313 submissions from 2009 to 2026 (58,892 after cleaning), we demonstrate three analytical approaches: (1) time-series analysis to identify temporal shifts in discourse volume, (2) dictionary- and rule-based extraction, together with Named Entity Recognition (NER), to recover origin–destination pairs and sociodemographic attributes from the structured titles, and (3) large language model (LLM)-based zero-shot classification, used only for classifying migration motivations. Validation against human-coded labels across three LLMs (GPT-4.1, Claude-3.5-Sonnet, and DeepSeek-V3) showed that GPT-4.1 achieved the highest agreement (mean κ = 0.566), with substantial agreement on push factors, moderate agreement on pull and enabling factors, and fair agreement on constraining factors. Applying this approach to 56,242 international migration posts, we found that enabling factors (individual-level resources) appeared most frequently (77.3%), followed by constraining factors (49.3%), pull factors (37.2%), and push factors (23.6%). For each method, we provide practical guidance on implementation, data access, and validation, and reflect on methodological limitations and ethical considerations.

More from our Archive