Rising AI Exam Scores Prompt Universities to Overhaul Assessments
A recent study by researchers at the University of Wollongong reveals that generative AI models have rapidly advanced to the point of outperforming the majority of undergraduate law students on final examinations. This marks a decisive shift from 2023 testing, when AI struggled to match human performance on intellectually demanding legal tasks. The new assessment evaluated nine models from five providers across compulsory criminal law and torts courses, utilizing internet access and enhanced reasoning features while strictly excluding course-specific materials. Responses were graded through a hybrid system involving subject coordinators and blind-reviewed tutors. The results demonstrate substantial improvement. AI-generated criminal law papers averaged 76.3 percent, surpassing 82.5 percent of the student cohort, while torts responses averaged 66 percent, outperforming 61 percent of peers. Notably, the critical analytical weaknesses that previously hindered AI performance on hypothetical legal scenarios have largely diminished. AI also significantly outperformed human students on essay components. However, the models exhibit jagged capabilities rather than uniform expertise. Performance fluctuates considerably between different providers and subjects, with some excelling in specific areas while failing in others. Additionally, reduced hallucination rates in certain models correlate with lower citation volumes, and source selection remains inconsistent. These advancements raise serious questions regarding academic integrity, particularly for unsupervised or take-home assessments. While AI has not yet achieved the reliability of a qualified legal practitioner, its current proficiency enables sophisticated academic misconduct and threatens the validity of traditional grading models. The research underscores that higher education must adapt to a landscape where AI integration is unavoidable yet requires rigorous oversight. In response, the study outlines a three-tiered assessment strategy to balance academic rigor with technological adaptation. First, foundational competency must be preserved through supervised, AI-free examinations, oral defenses, and in-class problem-solving, particularly in introductory courses. Second, universities should implement process-oriented evaluations that grade students on their AI collaboration techniques, including prompt engineering, source verification, error identification, and iterative refinement. Third, a relay assessment framework should mirror professional workflows. This model alternates between supervised independent drafting and AI-assisted revision, requiring students to critically audit, verify, and improve machine-generated content under strict conditions. Each stage is graded separately to measure both autonomous knowledge and AI supervision skills. Researchers caution that findings are limited to two law disciplines at a single Australian institution, and outcomes may vary across other fields, prompt configurations, or assessment formats. Subjective grading variables also introduce inherent testing limitations. Nevertheless, the study concludes that future graduates must possess strong foundational expertise to effectively audit, challenge, and enhance AI outputs. Educational institutions must transition from prohibition to structured integration, ensuring students can independently apply critical reasoning while mastering the supervision of increasingly capable generative systems.
