Increasing the Efficiency of Comparative Judgment for Writing Assessment: Warm‐Starting Estimation and Selection Using NLP
Michiel De Vrindt, Renske Bouwer, Anaïs Tack, Marije Lesterhuis, Wim Van den NoortgateAbstract
Comparative judgment (CJ) is an assessment method in which assessors compare pairs of essays and judge which is of higher quality. While CJ produces valid and reliable results, it often requires many judgments before quality scores become reliable. In this study, we examine whether natural language processing (NLP) can improve the efficiency of CJ by warm‐starting the estimation of quality scores and the selection of essay pairs. We predict essay quality by fine‐tuning a transformer model and use these predictions to construct informative priors in a Bayesian Bradley‐Terry‐Luce model. We then introduce three selection rules based on predicted scores and essay embeddings. In this study, we focus on Dutch writing assessments in secondary and higher education in two settings: one‐off assessments and recurrent assessments. Warm‐start estimation reduced the number of judgments required to reach a reliability of .70 by approximately 40%‐68% compared with likelihood‐based cold‐start estimation and by approximately 24%‐64% compared with Bayesian cold‐start estimation. To reach a reliability of .90, warm‐start estimation yielded smaller reductions of around 15% compared with cold‐start estimation. The warm‐start selection rules we constructed yielded smaller additional gains. In practice, a reliability of .70 was reached after two to four judgments per essay, and a reliability of .90 after seven to fourteen judgments. Overall, warm‐starting CJ substantially reduces assessor workload and improves the practical applicability of CJ as a summative assessment method.