DOI: 10.1177/21582440261476342 ISSN: 2158-2440
Generative AI in Writing Assessment: Comparing GPT-4 and Human Raters in Secondary Education
Wenting Bao, Miaomiao Zhang, Tianshu Wang, Roubing Li
This study explored the potential of artificial intelligence (AI) in Chinese writing assessment for secondary school students, and systematically analyzed the feasibility and limitations of AI-assisted educational assessment by comparing the scoring performance of the GPT-4 with that of human raters. The study used a randomized block group design to recruit 60 first-year middle school students to participate in the writing test (30 male and 30 female), and invited three language teachers with extensive teaching experience and the GPT-4 to score the test. The results of the study showed that (1) the AI scoring system showed high overall agreement with human raters (
ICC
= .84, 95% CI [.79, .89]), especially reaching the highest agreement in high-level essay assessment (
r
= .92,
p
< .001); (2) the AI system had a significant advantage in scoring efficiency, with an average scoring time of only 0.8 seconds/post and maintained a stable scoring performance (
CV
= 1.2%); (3) in the assessment of innovative expressions, the AI system scores (
M
= 76.5,
SD
= 5.8) were significantly lower than those of human raters (
M
= 82.3,
SD
= 6.1), reflecting its limitations in dealing with unconventional expressions. The study suggests a tiered assessment mechanism whereby the AI system is used primarily for basic, standardized scoring tasks, while reserving scoring tasks that require deep understanding and creative judgement for human raters. This study provides empirical evidence for the rational application of AI scoring systems in educational practice.