Artificial intelligence tools typically award higher marks than human graders and cannot be relied on to give an accurate indication of a student's performance, according to a new study published in the journal Assessment & Evaluation in Higher Education.
Researchers uploaded 50 undergraduate bioscience essays to two versions of ChatGPT and asked the AI to grade them against seven assessment criteria under four different prompting conditions. They found "significant discrepancies" between AI-marked essays and those graded by humans.
In all but one case, the AI models returned higher average marks than humans. In one instance, the difference between an AI grade and a human grade was 40 points on an essay where the top score was 100.
AI inflates low scores, deflates high ones
Lower-scoring essays tended to receive inflated marks, while higher-scoring essays received lower marks compared with human assessment. The AI was more aligned with human markers on mid-level essays.
The study said the large language models tested "varied considerably" in the marks they awarded and "were inadequate predictors of the human mark awarded to essays."
"Despite being relatively consistent at producing similar marks when using the same prompt on the same essay, when used to mark a range of essays of differing standards, there was poor alignment between the LLM-provided marks and the marks assigned by the original human marker," the paper said.
Why researchers tested AI grading
The research was primarily motivated to evaluate whether generative AI could be used as a formative benchmarking tool for students, in addition to feedback from academics, rather than whether it could replace human markers.
William Kay, senior lecturer in statistics at the University of Edinburgh and co-author of the report, said he had never questioned that responsibility for marking should remain a human one, but the study underlined that fact.
"At present, GenAI is unable to reliably assign a mark to a subjective piece of written work comparable to that of humans-even with extensive training of the LLM," he said.
Kay noted that universities are exploring whether AI's pattern-recognition capabilities could make marking more efficient and relieve pressure on staff. "The findings of this research indicate that at present this is not advisable. Aside from the ethical issues of submitting student work to GenAI tools without express consent, LLMs cannot, and should not, be relied upon for assigning grades to extended written work by students."
He added that while more sophisticated models may eventually improve their ability to mimic human judgement, "aligning marks between humans and GenAI may be hard to achieve."
The paper suggests LLMs might better predict "extreme" marks if grading criteria use "obviously distinct descriptions for each mark bracket," rather than terms like "good, excellent, outstanding," which "may be challenging for LLMs to differentiate."
Why this matters for educators
For teachers and academics weighing whether to use AI as a grading assistant, the study offers a clear caution: current tools cannot reproduce human judgement on subjective written work. The findings also carry implications for AI for Education more broadly, suggesting that pattern-recognition capabilities do not translate into reliable assessment.
Educators exploring these tools may find value in AI for Teachers training that focuses on appropriate use cases, such as drafting feedback or generating practice questions, rather than assigning grades. The study also notes that high levels of variation among human markers could make future alignment between humans and AI difficult to achieve, even as the technology improves.
Your membership also unlocks: