SCORECRAFT: A LARGE LANGUAGE MODEL-BASED SYSTEM FOR AUTOMATED EXAMINATION MARKING AND FEEDBACK GENERATION
Keywords:
Automated marking, Large language models, Automated essay scoring, Examination assessment, Educational technologyAbstract
Manual marking of examination scripts is time-consuming, prone to inconsistency, and often delays the delivery of feedback to students, particularly in large university classes. This paper presents ScoreCraft, a large language model (LLM)-based examination marking system designed to assist lecturers in evaluating student responses. The system allows lecturers to upload structured marking guides and student answer scripts, after which it generates draft scores and explanatory feedback for lecturer review and approval. The proposed system integrates marking guide ingestion, student response alignment, LLM-based scoring, feedback generation, and a human-in-the-loop moderation interface. By retaining the lecturer as the final authority over released scores, ScoreCraft is designed to support transparency, consistency, fairness, and accountability while reducing marking workload and turnaround time. The paper also presents an evaluation protocol based on agreement between system-generated scores and lecturer-assigned scores, using measures such as quadratic weighted kappa and mean absolute error. The proposed system provides a foundation for efficient and scalable examination assessment, with future extensions including multilingual assessment, handwriting recognition, and improved adaptation to diverse examination formats.