Large Language Models for L2 Italian Writing Assessment: Effects of Prompt, Fine-Tuning, and Data Balance
Wenqian Huang, Pengzhan Yang, Guangyuan YaoThis study systematically evaluates generative pre-trained transformer (GPT)-based large language models for the automated writing evaluation of Italian as a second language (L2). Drawing on 1,832 learner texts, we compare six GPT-based experimental conditions: zero-shot and few-shot prompting with GPT-4.1 and GPT-5, and two fine-tuned GPT-4.1 models trained on either class-balanced or proportionally distributed datasets. Model performance against human common European framework of reference for languages (CEFR)-based ratings is examined using exact agreement, correlation coefficients, Quadratic Weighted Kappa, precision, recall, and F1-scores. Results reveal a clear hierarchy: the balanced fine-tuned GPT-4.1 model achieves near-operational reliability, whereas all prompting-based methods, including GPT-5, remain substantially weaker. Training data balance emerges as critical, as class-imbalanced fine-tuning was associated with weaker performance on high- and low-level texts. The findings highlight the need for balanced, language-specific corpora and task-aligned fine-tuning in L2 Italian automated writing assessment.