A lecturer’s grade and an AI’s assessment of the same student work differed by as many as 40 points out of 100, according to recent findings that raise questions about the reliability of artificial intelligence in education. Researchers analyzed 85 studies, including a meta-analysis of 49 studies comprising 59 effects, to examine AI’s impact on STEM learning, but ultimately concluded that a reliable overall effect of generative AI use on learning could not be confirmed.
The team also discovered hidden grading instructions within the initial feedback generated by an AI system, highlighting a lack of transparency, Meta-Analysis says. “Checking and explaining AI-generated responses, using them transparently and taking responsibility for them are essential,” the Center for the use of ICT in pedagogical process reports.
AI Grading Varied From Human Scores By Up To 40 Points
Assessment scores produced by artificial intelligence deviated from those assigned by human lecturers by as much as 40 points on a 100-point scale, according to a systematic review of 85 studies. This discrepancy highlights a critical unreliability in individual grading, even when the overall average score for a group appears reasonable; personalization, often touted as a benefit of AI in education, is undermined by these inconsistencies.
Dr. Deepshikha at Queen Mary University of London observed a similar divergence between assessments completed with and without the assistance of AI, suggesting the issue is not isolated to specific systems or contexts. AI models excel at identifying surface-level qualities in student work, such as fluent writing and structural components, but struggle with nuanced evaluation of argumentation, evidence weighting, and the significance of errors within a specific learning environment.
The quality of a marking rubric proved to be the strongest predictor of agreement between human and AI graders; when a well-defined rubric was employed, AI performance improved and more closely mirrored human assessment. Deepshikha views the AI-generated assessment as potentially helpful commentary for students seeking rapid feedback, but emphasizes a fundamental difference between information provided by AI and truly effective feedback. She explains that high-quality feedback necessitates dialogue, meaningful connection to the student’s understanding, and careful, thoughtful formulation grounded in the student-teacher relationship.
Beyond simple score discrepancies, the review revealed instances of AI systems embedding hidden grading instructions within initial feedback provided to students, raising concerns about transparency and potential bias in automated assessment. A separate detector, while capable of flagging text as AI-generated, failed to provide any explanation for its decision, further obscuring the reasoning behind the assessment.
These findings reinforce the need for careful oversight and critical evaluation of AI-driven tools in education, as simply deploying the technology does not guarantee improved learning outcomes. Human judgement must remain paramount in educational assessment, not because humans are infallible, but because of our capacity to consider context, articulate reasoning, respond to student input, and accept responsibility for our decisions. AI can be a warning signal or a secondary opinion, but it cannot, and should not, be entrusted with making final judgements.
Meta-Analysis Reveals Limited Learning Gains From Generative AI Use
Researchers systematically reviewed 85 studies, focusing on 49 that provided quantifiable data on cognitive outcomes, and found no conclusive evidence that using AI reliably improves student learning. Observed benefits were more pronounced in knowledge acquisition than in the development of practical skills, and occurred primarily when AI assisted students rather than completing tasks for them. The meta-analysis revealed that successful integration of AI hinges on how it’s used, not simply that it’s used; what AI takes over in the learning process, and what the student does independently, are critical factors, according to the company.
Users frequently struggled with crafting effective prompts for AI systems, experienced cognitive overload, and demonstrated a tendency to accept inaccurate information without critical evaluation. These difficulties were particularly evident in STEM subjects when tasks involved verifying equations, writing code, or interpreting probability, suggesting that AI’s assistance requires careful oversight and validation.
The authors concluded that established methods for enhancing learning with AI remain elusive, despite the growing prevalence of these tools. Beyond the challenges of effective implementation, the review uncovered concerning issues with the transparency of AI-driven assessment. One AI system was found to embed hidden grading instructions within its initial feedback on student submissions, raising questions about potential bias and the need for clear, auditable evaluation criteria, the company says.
Practical work remains important for assessment, as it requires students to demonstrate abilities developed through sustained effort. Training in prompt engineering, critical evaluation of AI responses, and independent reasoning are valuable skills to cultivate, particularly for learners with less prior knowledge.
Cognitive Resilience & Practical Work Validate Traditional Assessments
While personalization is often touted as a benefit of AI-driven assessment, the research demonstrates that AI frequently misjudges the quality of submissions, overvaluing weaker work and undervaluing strong performances, despite appearing plausible when averaged across a group. This disparity highlights a critical issue; the ability to accurately gauge demonstrated competence, a cornerstone of higher education qualifications, is compromised when assessment relies on systems prone to such inconsistencies.
The systematic review, encompassing a meta-analysis of 49 studies comprising 59 effects, also identified specific challenges students face when interacting with AI tools, extending beyond simple inaccuracies. Learners struggle with crafting effective prompts to elicit useful responses, experience cognitive overload from managing the interaction, and exhibit a tendency to accept AI-generated information uncritically, potentially leading to poor decision-making regarding task delegation.
These difficulties underscore the need for a shift in pedagogical focus, prioritizing cognitive resilience, the capacity to maintain reasoning and judgement amidst constant access to AI, as a core skill for students. The research suggests that practical work, demanding active cognitive engagement, is essential for accurately assessing acquired knowledge and skills. The review also points to the importance of focusing on the process behind an answer, rather than simply the superficial characteristics of the output, advocating for assessment methods that evaluate judgement, production methods, and authorial responsibility.
Instead of solely relying on AI-content detectors, which have proven unreliable, the authors suggest a return to assignments that require students to analyze real-world processes and gather their own evidence, Meta-Analysis reports. In this scenario, AI can assist with clarifying instructions, formulating analytical questions, and improving clarity, but it cannot fabricate essential elements like observations, measurements, or process diagrams.
Rubric Quality Predicts AI-Human Agreement in Essay Grading
The consistency with which artificial intelligence aligns with human graders in evaluating student essays hinges significantly on the clarity and detail embedded within the assessment rubrics used, according to findings from Queen Mary University of London. This suggests that the structure defining expectations for student work is more influential than the assessment method itself in achieving consistent grading.
Deepshikha explains that AI tools perform optimally when paired with a rubric that explicitly outlines these criteria, transforming the AI’s output into a valuable commentary on student work rather than a definitive judgment of its merit. She emphasizes a fundamental distinction, stating, “feedback produced by AI tools is merely information or commentary, whereas high-quality feedback should also involve dialogue, be made meaningful and be formulated carefully, thoughtfully and with relationships in mind.” This perspective underscores the importance of human oversight in interpreting and contextualizing AI-generated assessments.
The need for robust rubrics extends beyond simply improving AI-human agreement; it also addresses concerns about the potential for manipulation within AI assessment systems. A recent security study revealed that researchers could influence assessment outcomes by subtly embedding instructions within student submissions, highlighting a vulnerability in AI-based evaluation. The study’s authors advocate for clear guidelines, enhanced educator expertise, and standardized security testing to fortify assessment robustness.
This proactive approach to rubric design is seen as a means of minimizing the impact of hidden prompts or instructions on the final grade. This finding reinforces the idea that a robust rubric, by clearly delineating expectations and assessment criteria, can help ensure that AI supports, rather than supplants, critical student thinking.
The authors concluded that no reliable ways of using AI to improve learning have yet been established, but that training in prompt writing, checking AI responses, and comparing them with one’s own reasoning is valuable. Institutional responses to these findings are already taking shape. At Queen Mary University of London, these reflections have prompted concrete steps, including investment in marking-rubric design, the introduction of formative assessment, and the establishment of institutional data governance.
These initiatives reflect a growing recognition that AI assessment is not a plug-and-play solution but rather a complex undertaking that requires careful planning, ongoing evaluation, and a commitment to maintaining human pedagogical judgment at the core of the evaluation process. The focus on rubric quality, therefore, represents a practical and actionable strategy for harnessing the potential of AI while mitigating its inherent risks.
Hidden Prompts Can Manipulate AI-Based Assessment Results
AI assessment systems, despite potential benefits, exhibit vulnerabilities to manipulation through subtly embedded instructions, a recent security study revealed. Researchers demonstrated the ability to influence grading outcomes by adding hidden prompts within seemingly standard student submissions, highlighting a critical flaw in current AI evaluation methods. This manipulation wasn’t about altering content; rather, it involved embedding instructions that steered the AI’s interpretation of the work, raising concerns about the integrity of automated grading.
The extent of potential grade discrepancies is substantial; comparisons between AI-generated scores and those assigned by human lecturers revealed differences of up to 40 points for the same student essays. While average scores across a group might appear reasonable, individual grades proved unreliable, undermining the promise of personalized assessment that AI is intended to deliver. However, even a robust rubric cannot fully mitigate the risk of hidden prompts, as the security study demonstrated.
The researchers did not use real student work, focusing instead on the system’s susceptibility to external influence, the company’s account states. The implications extend beyond simple grade inflation or deflation; the ability to manipulate assessment results raises questions about fairness, equity, and the validity of AI-driven evaluation in higher education.




See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.
