DOI: 10.3390/app16189268 ISSN: 2076-3417

From Black-Box Grading to Pedagogically Aligned AI Assessment: A Hybrid LLM–RAG Framework for Explainable and Scalable Automated Code Evaluation

Pablo Manuel Vigara Gallego, Ascensión López-Vargas, Ángel García-Beltrán, Javier Rodríguez-Vidal

Automated assessment of programming assignments remains a major challenge in higher education, particularly in large-scale courses where timely, consistent, and pedagogically meaningful feedback is difficult to provide, while Large Language Models (LLMs) have shown strong capabilities in code understanding and feedback generation, their use as standalone evaluators is fundamentally limited by inconsistency, lack of transparency, and weak alignment with instructional objectives. This paper argues that these limitations are not intrinsic to LLMs, but rather arise from their deployment as isolated components. In response, we propose a system-centric approach to AI-assisted assessment, introducing a hybrid framework that integrates LLMs within a structured, context-aware, and pedagogically aligned evaluation pipeline. The framework combines (i) explicit rubric-based decomposition of evaluation criteria, (ii) pedagogically guided prompting, and (iii) Retrieval-Augmented Generation (RAG) grounded in course-specific materials. Together, these components transform the evaluation process from a black-box prediction task into a traceable and reproducible decision process. The proposed approach is implemented in a real-world educational platform, EvaluaTeC, and evaluated on a dataset of 1287 programming submissions from 429 students. Experimental results show that the hybrid framework improves agreement with consolidated instructor reference grades (r=0.9059 vs. 0.7207 baseline), reduces evaluation error (MAE = 0.5134), and exhibited lower output variability in the recorded aggregate statistics, while maintaining practical latency and cost. Beyond numerical improvements, the system approximates key statistical properties of human grading and generates structured, pedagogically aligned feedback. These findings demonstrate that reliable AI-assisted assessment emerges from the integration of LLMs within structured and context-aware systems, rather than from model capabilities alone. This work contributes a principled framework for explainable and scalable automated assessment, advancing the design of trustworthy AI systems in education. This shift reframes automated assessment as a systems problem rather than a purely model-centric task.