Article

Google’s Patent-Lawyer Study: Measure AI Output and Human Judgment Separately

Google’s patent-lawyer study separates AI-assisted drafting gains from unaided judgment. Here is how to evaluate output, efficiency and professional learning.

Editorial illustration for Google’s Patent-Lawyer Study: Measure AI Output and Human Judgment Separately: a document represents the research briefing. Not documentary evidence.

On October 7, 2026, Google Research published an explanation of a three-month patent-lawyer experiment: AI access improved assisted drafting, but the later advantage on an unaided review task was concentrated among experienced lawyers. For professional-services leaders, the practical question is whether an AI rollout improves delivered work, develops the people doing it, or both.

The underlying NBER working paper was issued in September 2026; these are not results first released today. Its useful lesson for evaluation is straightforward: a stronger AI-assisted draft does not establish stronger unassisted judgment.

What the trial measured

Researchers randomized 133 patent lawyers at 11 US intellectual-property firms: 90 received early access to Google Labs’ prerelease InFlow assistant and 43 entered the control group. Both groups were offered AI-basics training; the treatment group also received tool-specific training. The intervention therefore combined access and training, rather than isolating a model’s capabilities. The paper’s methods describe staggered cohorts between May 2025 and February 2026, with roughly three months per cohort.

Lawyers completed patent-drafting exercises using hypothetical invention materials. For the final assessment, participants were instructed to correct an existing flawed draft without AI—a different task, called redlining. Third-party patent attorneys, blinded to assignment, rated quality across five dimensions, including technical accuracy and clarity. These were expert-scored exercises, not patent-office decisions or court outcomes. Methods and grading rubric.

The authors’ preferred expert-rating estimates are summarized below; completed samples differ by task. Tables 1, 3, 4 and 6 provide the denominators and estimates.

Assessment

Lawyers completing it

Reported treatment effect versus control

Early assisted drafting, approximately day 10

105

+0.34 standard deviations; p = 0.03

Later assisted drafting, approximately day 90

98

+0.38 standard deviations; p = 0.01

Unaided redlining at the final assessment

91

+0.32 standard deviations; p = 0.04

Standard deviations describe differences relative to the spread of control-group scores; +0.38 SD does not mean 38% better work. Nor can the drafting and redlining effects be divided to calculate a skill-retention rate: the tasks, outcome distributions and completed samples differ.

The overall average misses the training question

On unaided redlining, the senior-lawyer estimate was +0.45 SD (p = 0.02), while the junior estimate was −0.03 SD (p = 0.89). The latter is near zero and statistically inconclusive, not evidence that junior lawyers lost skill. The analysis defined juniors as having fewer than seven years of experience and seniors as having seven or more. Table 6 and the experience definitions.

Junior lawyers had larger estimated benefits on assisted drafting, although the junior–senior differences in treatment effects were not statistically conclusive. Their unaided scores were more mixed, with more poor and more good scores rather than a clear average advantage. Those group distributions do not identify which individuals improved or declined. Drafting and redlining results.

For a rollout owner, this supports two separate decisions. An assistant may earn continued use because it improves output, while the organization still needs evidence that its apprenticeship and review practices develop independent judgment. A positive company-wide average is not enough to settle the training question for less-experienced staff.

Speed needs its own accounting. The authors report about ten minutes saved on the early drafting task against an approximately 112-minute control baseline. A similar later estimate was statistically uncertain. Timing was self-reported, and exploratory workplace measures did not establish firm-wide efficiency gains. These results cannot be translated directly into lower client bills or staffing savings. Section 5 and Table 8.

What this cannot establish about learning

The paper’s methods and discussion leave several important limits. There was no baseline skills test, so interpreting endpoint differences as learning relies on initial balance from randomization rather than observed individual before-and-after growth. Only 91 of 133 randomized lawyers supplied the final redlining outcome; one cohort finished before that task was introduced, so the missing outcomes are not all ordinary dropout.

Researchers could not directly enforce the no-AI instructions. The authors report sensitivity checks, but those do not establish perfect adherence. Three months also cannot demonstrate career-long skill development or persistence long after access ends. The sample came from firms already doing sophisticated patent work for Google; transfer to other professions remains an open question. Study limitations.

It would also be premature to conclude that reducing effort necessarily prevents learning. In a separate set of writing experiments by Benjamin Lira and colleagues, participants who practiced with AI later wrote better cover letters without it despite exerting less practice effort. That is not a replication in patent law; it is a reason to test learning in the actual task and workflow rather than assume a universal effect.

A pilot that evaluates the work and the worker

The following is an evidence-informed design proposal. RohitAI has not run or validated this pilot. It assumes the job still requires independent judgment and that representative, consistently graded exercises can assess it.

Decision

Evidence to collect

What it establishes

Is the assistant useful for production?

Blind domain-quality ratings, including critical errors, on comparable assisted and usual-workflow tasks.

Whether delivered work improves with assistance.

Is the full workflow more efficient?

Elapsed time, tool cost, reviewer time and rework across all attempts through acceptance.

Whether gains survive verification and correction costs.

Is independent capability developing?

Equivalent unseen no-AI tasks before and after exposure, compared with a control where feasible.

Whether unaided judgment changes, separately from assisted output.

  1. Define the unaided task before rollout. Use permissioned or synthetic materials and qualified domain reviewers. Score whether participants actually correct important defects and can explain their changes, not just whether the prose looks polished.

  2. Keep the comparison interpretable. Use comparable groups or a staggered rollout where feasible, blind graders to assignment, and record tool versions, training and assistance rules. Choose experience bands that fit the organization rather than importing seven years as a universal threshold.

  3. Report completion rates and uncertainty by baseline expertise. Retain lower-scoring outcomes as well as averages. If the sample cannot distinguish a meaningful change from noise, report that uncertainty rather than declaring no effect.

One training variant worth testing is to require an initial independent assessment before revealing AI suggestions, followed by a reason for accepting or rejecting important edits. Compare it with ordinary assistance. The patent experiment did not independently randomize this interaction pattern, so it does not prove that the variant preserves expertise.

This complements the contracting-agent evaluation guide, which focuses on whether an automated workflow behaves correctly. Here the additional question is whether the person supervising that workflow can find and repair mistakes independently.

If assisted quality improves but unaided performance does not show a reliable gain, record a production benefit without a demonstrated training benefit. If independent performance clearly worsens, investigate task design and support before expanding dependence. Keeping those decisions separate lets a team retain useful assistance without mistaking its output for evidence of human learning.

Methodology: AI-assisted reporting and analysis of the published paper, Google’s explanation and supporting primary research, checked October 7, 2026. Results are author-reported; no hands-on tool evaluation or independent replication was performed. This is a non-peer-reviewed working paper. NBER’s disclosures state that Google funded the experiment’s direct costs and identify co-author employment, equity and contractor relationships.