Automated Scoring of Long‐Context Essays Using Linear‐Complexity Language Models
Christopher OrmerodAbstract
Automated essay scoring (AES) increasingly uses pretrained language models, but many operational systems still impose fixed input limits that truncate essays and may weaken score validity. We study how truncation and model architecture affect scoring for long‐form writing using more than 20,000 responses. We compare a modern Transformer encoder, ModernBERT, with a linear‐time state‐space model, Mamba‐130M, which supports efficient long‐context processing. Evaluation follows the standard framework for the evaluation of automated scoring using Quadratic Weighted Kappa, exact agreement, and standardized mean difference. A paired permutation test (1,000,000 resamples) confirms that truncation produces a statistically significant reduction in QWK for both architectures (), although the effect size is modest. To assess truncation‐related bias, we also examine Kullback‐Leibler divergence on propensity score‐matched samples and Jeffrey's divergence. Our results show that truncation harms not only scoring accuracy but also construct validity by forcing models to judge incomplete evidence of writing quality. By contrast, the state‐space model shows favorable distributional behavior and much lower inference latency as sequence length grows, while maintaining competitive scoring performance. Overall, the findings highlight an important trade‐off between efficiency and validity in AES and suggest that linear‐time architectures are a promising basis for more faithful evaluation of long‐form writing.