DOI: 10.1177/00208523261473520 ISSN: 0020-8523

Web crawling in public administration research

Jakob Marquardt, Jonathan Gerber, Claire Kaiser, Jana Machljankin, Reto Steiner

Improvements in computational power and artificial intelligence technology are transforming science, including public administration research. Techniques from the computer sciences, such as web crawling, can improve data collection and analysis. However, they remain underutilized, and the literature exploring their potential in public administration research is limited. This paper aims to fill this gap by analyzing the opportunities and limitations of web crawling for data collection by answering two research questions: what opportunities and limitations does web crawling offer compared to manual web data collection, and under what conditions is web crawling an effective data collection method in public administration research? We draw on theoretical considerations from existing literature, prior empirical evidence, and insights from our own research. Based on our analysis, we highlight the benefits of the immense scalability, efficiency, and consistency that a web crawler can provide for creating large datasets, while cautioning against their use in smaller-scale research, where high setup costs, expertise requirements, and limited flexibility may outweigh potential advantages. Additionally, we provide guidelines on when and how to use web crawlers usefully in public administration research.

Points for practitioners

The paper evaluates the opportunities and limitations of using web crawling as a data collection technique in public administration research. The evaluation is based on existing empirical evidence, as well as experience from a research project.

Web crawling can be a highly useful tool that complements traditional qualitative and quantitative methods in various contexts.

When comparing web crawling to the manual collection of web data, two points stand out as particularly relevant, both within and beyond research contexts: first, web crawling generally comes with significant setup costs but is highly efficient in execution. This makes it highly scalable and particularly suitable for collecting large amounts of data. Second, proper implementation of a web crawler requires significant information technology expertise. Close collaboration between those who want to use the data and the information technology experts implementing the data collection is recommended to ensure that relevant, high-quality data can be obtained.

More from our Archive