Can AI Grade Exams Better Than a Law Professor?

A recently published paper co-authored by Eric Posner explores what large language models reveal about fairness, bias, and judgment in legal education grading
Robot professor grading law exams
AI-generated image
A graphic that says Scholarly Pursuits

Editor’s Note: This story is part of an occasional series on research projects currently in the works at the Law School.

For law professors, grading law school exams is among the most consequential—and least beloved—parts of their job as legal educators. It is time-consuming, cognitively demanding, and fraught with concerns about fairness and consistency. 

In a new paper he coauthored, Eric Posner, the Kirkland & Ellis Distinguished Service Professor of Law, asks a provocative question: Can artificial intelligence grade law school exams as well as humans do—and, if so, what might that tell us about how grading really works?

The project grew out of Posner’s own teaching experience. After grading a large Contracts exam, he and a research assistant ran the same exams through a large language model months later, after grades had already been released to students. What they found was eye-opening.

“There was a very high correlation between the grades that I gave and the grades that ChatGPT gave,” Posner said. “And it was really surprising because I didn’t try very hard to give guidance to the LLM. It just did it by itself.”

That initial experiment led Posner to join a team of five professors from different law schools and disciplines to conduct a more systematic study using multiple courses, instructors, and grading approaches. The study utilized final exams from four traditional law school core subjects: civil procedure, contracts, torts, and corporations. The AI was asked to regrade the exams human professors had graded. In some cases, it was given little guidance, and in other cases detailed grading instructions.

Across those settings, the researchers found consistently high correlations between human grading and AI grading—sometimes approaching the level of agreement seen when the same professor re-grades their own exams after time has passed.

Human grading is imperfect—by design

One of the paper’s core contributions is reframing how legal educators think about grading accuracy. Rather than treating human judgment as a gold standard, Posner argues that grading has always involved a degree of subjectivity and noise.

“There isn’t an objective grade,” he explained. “Probably for every exam, there’s a distribution of reasonable grades.”

Human graders, Posner notes, are vulnerable to well-documented cognitive effects—fatigue, drift, compression, and inconsistency—that arise precisely because grading is demanding and repetitive.

“Anybody who’s graded has experienced these problems,” he said. “You realize that if you grade a batch of exams, the numbers you give for the first exam are going to influence your subsequent exams.”

AI systems, by contrast, do not tire, drift, or lose focus in the same way. That difference, the paper suggests, raises a deeper question: When human and machine graders disagree, who is actually making the mistake?

Bias, consistency, and legitimacy

The paper also engages longstanding concerns about bias and legitimacy in grading. While law schools typically rely on anonymous grading to reduce bias, Posner notes that concerns remain about writing style and other subtle cues.

“There’s always been concerns that the grader will have biases,” he said. “And again, there’s a question of whether the LLM is going to have the same bias, or a different bias, but that’s something one can actually test.”

Rather than compounding legitimacy concerns, Posner believes AI may ultimately reduce them.

“I suspect strongly that AI is going to be more consistent and therefore fair,” he said. “And in terms of legitimacy, it seems to me it would follow that the grades would be more legitimate if there was less human error.”

One of the study’s most striking findings is how closely AI grading tracks human grading when a detailed rubric is used. That result, Posner suggests, sheds light not only on AI performance, but also on how professors themselves grade.

“Although grading is so important, I don’t think law professors are ever trained on how to grade,” he said. “As a result, law professors grade exams in very different ways.”

Detailed rubrics can improve consistency between human and AI grading, but they also constrain judgment, especially when students offer creative or unexpected answers. AI systems, Posner notes, mirror this tradeoff: the more structured the rubric, the more closely the machine will approximate a particular professor’s approach.

A supplement, not a substitute

Despite the strong results, Posner does not believe law professors will be replaced by AI anytime soon. At least for now, he sees AI as a tool to assist, rather than supplant, human judgment.

One possibility the paper explores is using AI as a check on human grading—flagging exams where machine and professor assessments diverge sharply, prompting a second look.

The potential implications of the study extend beyond legal education. The same reasoning that applies to grading, Posner suggests, may also apply to judicial decision-making and legal practice, domains where ambiguity, judgment, and explanation are central. 

Posner, who has a burgeoning amount of scholarship exploring issues at the intersection of law and AI, believes active engagement with AI is critical for both law students and legal scholars. 

“It's important for law schools to educate their students on to how to use AI in legal practice because that is now a necessary skill,” he said. “I also think it's important for law professors to continue to do research on AI because it's going to be a big thing and it’s going to raise all kinds of questions.”