An average score of 96 percent on a Welfare Economics midterm at Brown University prompted professor Roberto Serrano to suspect widespread cheating, despite intentionally designing the exam to be more challenging than previous iterations. This spring, Serrano allowed his class of 86 students, a significant increase from the typical 30, to take the exam at home following a campus shooting, a decision he later regretted. “Historically the average grade in the midterm of this course has ranged between 65 and 80 percent,” Serrano said, noting the unusually high results. After confirming his suspicions with ChatGPT, which mirrored student responses with answers that were “kind of correct, but very off and with a very convoluted style,” Serrano switched to an in-person final, leading to a historic low average score of 48.6 percent and ultimately, 19 failing grades.
Take-Home Midterm and Suspected AI Use
This result represents a significant jump from the historical 65 to 80 percent range typically seen for the same exam, even though Serrano intentionally designed this iteration to be more challenging than previous versions because it was take-home. The shift in performance coincided with a substantial increase in class enrollment; 86 students participated this semester, compared to a typical 30, a rise Serrano attributes to the promised take-home exams. Serrano’s concerns arose from the unusually high scores and a noticeable stylistic consistency across student submissions. He and his graders subjected the midterm responses to analysis by ChatGPT, discovering striking parallels between the students’ work and the AI’s output.
One particular question, requiring a direct mathematical proof, was addressed by both ChatGPT and numerous students using a “contradiction argument,” which gave the right answer but was “very contrived” and which Serrano could tell wasn’t written by a human. Following these findings, Serrano altered the final exam to an in-person format, a decision that led to more than a dozen students dropping the course and a further increase in failures. Nine students did not even attempt the in-person final. The results confirmed his suspicions; three students received failing grades, and the average final exam score plummeted to 48.6 percent, a historic low, far below the previously consistent minimum of 65 percent.
Serrano subsequently declared the midterm void, weighting the final exam at 80 percent and previously set the passing grade at 50 percent or higher. “I am not declaring the midterm void for now. I am going to give the class a chance to prove me wrong,” he initially wrote to his students, anticipating the outcome. Despite his efforts to address the situation, Serrano reports a “meek” response from administrators, who requested individual complaints against each suspected student, a process he deems “ridiculous”.
They don’t get rewarded for dealing with a 60-person case of cheating,” she said.
Tricia Bertram Gallant, Director of the Academic Integrity Office and Triton Testing Center at the University of California, San Diego
Welfare Economics Course Score Discrepancies
The increasing prevalence of artificial intelligence tools presents challenges to academic integrity, prompting a re-evaluation of assessment methods and institutional responses. While concerns about plagiarism are longstanding, the sophistication of current AI raises the stakes, particularly in disciplines requiring analytical and argumentative skills. Roberto Serrano, an economics professor at Brown University, experienced this firsthand this semester when a significant shift in student performance on a Welfare Economics midterm prompted an investigation into potential AI-assisted cheating. Serrano’s class saw an increase in enrollment to 86 students, a figure he attributes to the promised take-home exams. The AI-generated answers were described as “kind of correct, but very off and with a very convoluted style.” The professor noted a specific example where ChatGPT, and a number of students, opted for a contradiction argument to prove a mathematical statement, a method Serrano deemed very contrived and indicative of non-human authorship.
The results of the in-person final, with an average score of 48.6 percent, a historic low compared to previous averages that never fell below 65 percent, confirmed his suspicions. Eighteen students dropped the course after the midterm, while nine did not attempt the final exam. Previously, he had set the passing grade at 50 percent or higher. Despite these measures, 19 students ultimately failed the course, highlighting the scale of the issue and the difficulties in balancing academic rigor with equitable assessment in an age of readily available AI tools.
Historically the average grade in the midterm of this course has ranged between 65 and 80 [percent], and this exam was harder than the exams I wrote in the past, because … take-home is an opportunity to challenge the class a little bit more, given that you’re giving the students unlimited time.
Eighteen students dropped the course after the midterm, while nine did not take the final exam. Serrano noted that only a few students maintained comparable performance levels between the two exams, solidifying his initial suspicions. Ultimately, 19 students failed the class. Serrano’s attempts to address the issue with Brown’s Standing Committee on the Academic Code initially met with limited response; the committee requested individual complaints and exam copies for each suspected case, a process Serrano deemed “ridiculous.” He expressed concern that the proposed investigation would rely on AI-detection tools known for inaccuracies. The situation highlights the challenges institutions face in balancing academic rigor with the evolving landscape of AI-enabled academic misconduct, and the need for proactive, adequately resourced responses to maintain academic integrity.
That is, if the distribution of the final exam is roughly similar to the distribution of the midterm, I will count the midterm. Otherwise, which is of course what I expect to happen, I will declare the midterm void and reweigh the final accordingly.
The surge in accessible artificial intelligence tools is forcing universities to confront a new era of academic dishonesty, with Brown University recently experiencing a particularly stark example of the challenges ahead. The results of the in-person final were telling; the average score plummeted to 48.6 percent, a historic low, demonstrating a clear divergence from previous performance. Brown spokesperson Brian Clark confirmed that the university treats all academic integrity allegations seriously and noted that Serrano had not yet provided the necessary details to initiate a formal investigation.
Brown treats every allegation of academic integrity with the utmost seriousness. In regard to this economics course, multiple academic leaders from Brown were in touch with the faculty member who raised concerns to provide details about how the allegations raised could be formally adjudicated. To date, the faculty member has not provided the necessary details to the Standing Committee on the Academic Code to pursue this path toward resolution.
See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.
