DOI: 10.3390/educsci16081227 ISSN: 2227-7102

Assessing Linguistic and Content Characteristics of First-Grade Students’ Texts Using Generative AI

Daniel Then, Caroline Theurer, David Cropley, Sanna Pohlmann-Rother

The present study examines how accurately a common large language model (ChatGPT 5.0) assesses linguistic and content characteristics of texts written by first-grade students. The focus is on the consistency of LLM assessments and their alignment with human assessments. The database consists of N = 539 texts written by elementary school students. The results show that, contrary to our expectations, LLM assessments of linguistic text characteristics are not generally more consistent than LLM assessments of content text characteristics. Furthermore, alignment between LLM and human assessments is higher for the assessment of single linguistic characteristics (such as the use of cohesive devices) than for the assessment of single content characteristics (such as originality), but not in general. Thus, LLM assessments of linguistic text characteristics in early elementary school are not generally more accurate than LLM assessments of content text characteristics. Based on these findings, we discuss implications for using LLMs as assessment tools in early stages of schooling.

More from our Archive