Grades collapsed from 96% to 48% when the professor banned AI from the exam
An economics professor at Brown University suspected most of his 86 students had used AI on a take-home exam when the average hit 96%, versus the usual 65-80%. Many had used a convoluted mathematical proof identical to ChatGPT’s instead of the more obvious direct approach.
He warned the students and gave the final exam in class, proctored. Eighteen students dropped the course, nine didn’t show up. The average crashed to 48.6%, the worst result ever recorded for that course. Only a few scored close to the first exam. Nineteen students failed.
Why this matters to you. Two recent studies confirm the same pattern at scale. A Chinese study on 26,000 students showed that six months after AI adoption, homework grades rose 18% while supervised exam grades fell 20%. The full effect emerges over about two years, with top students losing up to 24% of their performance. A UC Berkeley study on 500,000 grades shows that in courses with many unproctored assignments, the A percentage increased by 13 percentage points after ChatGPT’s launch.
If you teach or are learning to use AI, this is the core issue: AI that does the work for you doesn’t empower you, it hollows you out. The question isn’t about banning tools, it’s about understanding where foundational competence must hold up even without them.
In detail
The context: what came before
Take-home exams existed forever, but relied on an implicit academic honesty pact. ChatGPT and other models broke that pact because an average student can now get professional-quality answers in seconds, with no way for a professor to tell the difference by looking only at the final result.
Roberto Serrano, the Brown professor in this case, ran the exam questions through ChatGPT and got answers nearly identical to many students’, including an unusual mathematical proof choice that no student had ever used before. The 96% average was an obvious statistical anomaly: that course had always averaged between 65% and 80%.
The numbers from large-scale studies
The Chinese study tracked over 26,000 students from seventh through twelfth grade over 30 months. After six months of AI use:
- Homework grades rose 18%
- Completion time dropped from 64 to 45 minutes
- Supervised exam grades fell 20%
- On college entrance exams, long-term loss was between 18% and 24%
- 81% of frequent AI users showed the full pattern: quick homework, high homework grades, collapsed exam performance
Top students were hit hardest: they lost up to 24% of their performance. The full effect emerges over about two years, suggesting progressive skill deterioration.
The UC Berkeley study analyzed over 500,000 grades at a large research university in Texas. In courses with many writing or programming assignments, the A percentage rose 13 percentage points after ChatGPT’s November 2022 launch. The effect concentrated in unproctored assignments: courses relying heavily on take-home work saw a 16 percentage point greater increase than courses relying more on proctored exams.
What changes
The Brown case shows institutional response: the professor abolished the take-home exam and moved the final to 80% of the total grade. The university responded “timidly,” according to Serrano, asking him to report each case individually. He wanted stronger action: “We can’t afford to have a society where a significant fraction of our best young minds thinks cheating is OK. That leads to a declining society, a failed society… We can’t choose to become idiots.”
The limits of what we know
The studies show clear correlation, but can’t distinguish between students using AI as support versus students using it as complete replacement. The pattern emerges at aggregate level: it says nothing about how any single student is using the tool.
The Brown case rests on strong suspicion but not direct proof of every individual cheating case. The unusual proof and anomalous average are strong indicators, but not proof for each student.
Practical implications
If you teach, the numbers say unproctored homework no longer measures what you thought it did. Assessment design must change: more proctored exams, or assignments requiring you to show your process alongside results.
If you’re learning to use AI, the lesson is clear: AI that does the work for you doesn’t empower you, it weakens you. Using it on assignments you should solve alone leaves you with better grades today and worse skills tomorrow. The collapse comes when you have to work without a net.
If you’re building AI education tools, design matters: a tool that answers directly instead of guiding creates dependency, not learning.