A Stanford University-led study found that law professors preferred answers generated by large language models to responses written by fellow instructors in roughly 75% of blind comparisons. The 16 participating professors, representing 14 U.S. law schools, completed 2,918 anonymized head-to-head evaluations of short-answer tutoring for first-year contracts courses.
The models - Google Gemini 2.5 Pro and NotebookLM - recorded average win rates of 75.92% and 74.75% against the human instructors. The study focused on whether AI-generated legal explanations met the professional standards professors use when they judge arguments that involve uncertainty and competing interpretations, rather than measuring responses against a single correct answer.
AI answers lead across every question category
The 40 questions were drawn from a pool created by the participating contracts professors. Each instructor answered questions they had not written, while the AI systems produced responses to the same material. Answers were anonymized and lightly standardized to mask their origin. The questions spanned case and code recall, legal doctrine, hypothetical scenarios, and policy issues - including ones without a clearly correct answer.
Gemini 2.5 Pro and NotebookLM outperformed faculty-written answers in every category. The study focused on short-answer tutoring for first-year contracts, a domain that overlaps with many of the contract analysis tasks covered in AI Learning Path for Paralegals. Even for questions that required weighing competing arguments rather than identifying one factual response, the AI advantage held.
"We were frankly surprised by the magnitude of the results," said Julian Nyarko, a Stanford Law School professor and co-author of the paper. "These weren't just simple questions with obvious answers. Many of them required synthesizing complex material, applying it to new situations, and explaining legal concepts in ways that would help students develop their own analytical skills."
Fewer answers flagged as potentially harmful
The professors also rated whether individual responses could hinder student learning. AI-generated answers were flagged as harmful in 3.53% of cases, while faculty-written answers averaged 12.06%. The gap was consistent: Gemini was flagged 3.41% of the time and NotebookLM 3.64%, whereas individual instructors' harmful-answer rates ranged from 1% to 39.75%.
Answer length and other writing characteristics explained only part of the preference. When the researchers controlled for structure, clarity, confidence, legal references, and pedagogical support, the models still beat the rates those features alone would predict.
Study design tested professional judgment, not one right answer
Most past AI tutoring evaluations have used subjects with a fixed answer. Here, the researchers asked professors to choose the response they would rather give a student during office hours. Sarath Sanga, co-author and Yale Law School professor, said: "In most fields where AI gets tested, there's a right answer. In law, there often isn't. Two opposing arguments can both be good. What we wanted to know is whether AI can meet the latent professional standard that lawyers use to evaluate each other's arguments. In this case, the answer was yes."
Professors showed stronger agreement on quality than individual preferences alone would suggest. Their agreement was highest on policy questions, and NotebookLM's access to the casebook did not give it a meaningful edge over Gemini, even on questions whose answers were contained in the assigned material.
Researchers caution against wholesale adoption
The study involved a small, self-selected group: 16 of the 60 invited professors completed the work. They were more likely to be tenured and to come from top-14 law schools than the broader pool. The authors note that the shared professional standard identified may reflect these particular instructors, not legal educators generally.
The research was limited to brief written answers in a single course. It did not track longer tutoring conversations, student retention, actual academic performance, or whether regular AI access changes critical thinking. "Our study evaluates the quality of answers given by AI tools. But how to implement these tools to most effectively improve student learning is still an open question," Nyarko said. "The conversation should shift from whether AI can give accurate, high quality responses to how we can deploy it responsibly to the benefit of our students."
The authors propose course-based randomized controlled trials as a next step. They also recommend that any legal education system using AI include clear limits, citations, refusal mechanisms for uncertain questions, and escalation routes to human instructors.
Why this matters for legal professionals
For practicing lawyers, judges, and law school instructors, the study punctures the assumption that AI cannot yet match the nuanced judgment required in legal explanation. It does not suggest replacing professors or senior lawyers, but it does mean that blanket skepticism toward AI-generated legal reasoning is increasingly hard to defend. The real work ahead lies in determining where and how to deploy these tools - in tutoring, in first-draft contract review, in legal research synthesis - with clear guardrails. Professionals who build a working knowledge of these systems can evaluate them on their actual output rather than on intuition, a skill set that courses like AI for Legal Professionals Courses are designed to develop.
Your membership also unlocks: