DOI: 10.31681/jetol.1882892 ISSN: 2618-6586
Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language
Yunus Özdemir, Emine Aşçi This study investigates the effectiveness and prompt sensitivity of leading large language model (LLM)-based AI tools (ChatGPT, Gemini, Claude, Deepseek) in grading open-ended exam questions in an undergraduate history course (Atatürk’s Principles and History of Revolution), compared to human evaluators. A comprehensive dataset comprising 72 distinct open-ended responses (collected from 24 students) was scored by both the course instructor and seven AI models using five different prompts with varying levels of detail. These prompts ranged from a basic zero-shot instruction to progressively more structured designs: expected topic headings, a weighted criterion-based rubric, partial coverage with language proficiency, and relative (norm-referenced) scoring. Correlation and statistical analyses (paired samples t-test, Wilcoxon signed-rank test) revealed that while models showed low agreement with the human evaluator when using unstructured, basic prompts (zero-shot), they achieved high agreement (r > .80) when provided with structured prompts and explicit rubrics, particularly in the cases of Gemini and Claude. The findings highlight that for effective AI-based assessment, prompt design and the definition of criteria are more critical than model selection in mitigating "generosity bias." Generosity bias here denotes the models’ systematic tendency to assign higher scores than the human rater; the qualitative evidence indicates that it stems primarily from the models rewarding fluent, lengthy, and well-structured answers even when their factual content is incomplete, a tendency that explicit rubrics substantially reduced. Designed as an exploratory case study, these results provide significant empirical evidence regarding the optimization of AI as an assistive assessment tool, specifically within the context of Turkish, a morphologically rich language.
More from our Archive
-
DOI: 10.68381/jca02008 2026
Proximal Smoothness and the Lower-C
2
Property F. H. Clarke, R. J. Stern, P. R. Wolenski
-
DOI: 10.68381/jca13044 2026
Characterizations of Prox-Regular Sets in Uniformly Convex Banach Spaces Frédéric Bernard, Lionel Thibault, Nadia Zlateva
-
DOI: 10.68381/jca15047 2026
Brøndsted-Rockafellar Property and Maximality of Monotone Operators Representable by Convex Functions in Non-Reflexive Banach Spaces Maicon Marques Alves, Benar Fux Svaiter
-
DOI: 10.68381/jca16027 2026
Proximal Smoothness and the Exterior Sphere Condition Chadi Nour, Ron J. Stern, Jean Takche
-
DOI: 10.68381/jca16053 2026
A New Old Class of Maximal Monotone Operators Maicon Marques Alves, Benar Fux Svaiter
-
DOI: 10.68381/jca13045 2026
Maximal Monotonicity via Convex Analysis Jonathan Borwein
-
DOI: 10.68381/jca08009 2026
Variational Inequalities and Regularity Properties of Closed Sets in Hilbert Spaces Giovanni Colombo, Vladimir V. Goncharov
-
DOI: 10.68381/jca17060 2026
Existence and Uniqueness of Solutions for Non-Autonomous Complementarity Dynamical Systems Bernard Brogliato, Lionel Thibault
-
DOI: 10.68381/jca01001 2026
Variational Sum of Monotone Operators H. Attouch, J.-B. Baillon, M. Théra
-
DOI: 10.68381/jca22017 2026
Weak Convexity of Sets and Functions in a Banach Space Grigorii E. Ivanov