DOI: 10.1111/exsy.70381 ISSN: 0266-4720

A Topic‐Driven Mixture of Experts With Low‐Rank Adaptation for Multi‐Trait Scoring of Teachers' Reflective Essays

Jiacheng Gu, Gong Wang, Fei Jiang, Zhiying Wang, Haibo Peng

ABSTRACT

Current methods for multi‐trait essay scoring mainly focus on style‐based traits such as grammar and structure, with limited attention to multiple content‐level traits. This limitation prevents them from effectively assessing the fine‐grained content quality of essays, resulting in evaluation inaccuracies where high‐scoring essays may reflect formulaic writing rather than insightful content. In this paper, we propose a novel multi‐trait scoring task based on teachers' reflective essays, aiming to evaluate instructional insight conveyed through concrete content. To explore content‐oriented scoring of teaching reflection, we first construct a real‐world dataset, PTRE, consisting of 1820 reflective essays written by pre‐service teachers. Each essay is annotated with scores for four instructional content traits: Textual Expression, Problem Awareness, Comprehensive Thinking and Continuous Growth. Additionally, the dataset includes partial sentence‐level annotations highlighting trait‐relevant content as supporting evidence for each trait. We further design TMoLE, a topic‐driven mixture of experts framework for content‐oriented multi‐trait scoring. It learns trait‐specific topic distributions guided by seed words derived from sentence‐level annotations under expert guidance, ensuring semantic alignment with each scoring trait. Using these distributions as routing signals, TMoLE dynamically activates experts implemented with low‐rank adaptation to precisely capture trait representations and improve multi‐trait scoring. Experimental results on the constructed PTRE dataset show that TMoLE achieves an average Quadratic Weighted Kappa (QWK) of 0.6138. It achieves the highest average QWK among the compared methods and the best results on Textual Expression, Problem Awareness and Continuous Growth, supporting its effectiveness for content‐oriented multi‐trait scoring.

More from our Archive