GitHub announced ReviewBench’s public research preview on October 5, 2026, giving engineering teams and AI-reviewer builders a shared way to evaluate code-review agents. The release includes a public pull-request corpus, labeled findings, a leaderboard and a bring-your-own-agent evaluation path.
Use it to shortlist reviewers and compare configurations before a repository-specific pilot. The useful question is whether a reviewer catches the issues your team cares about at a comment volume it can handle—not simply which row ranks first.
Check corpus fit and what counts as a useful finding
ReviewBench’s reference findings combine human comments, issues inferred from follow-up changes, static analysis and model-generated reviews, followed by deduplication and auditing. These are judgments of usefulness under a rubric, not an exhaustive inventory of every defect. The methodology explains how that reference set is built.
The published corpus manifest contains 219 pull requests from 187 repositories, with 19 identified language labels and one unknown entry. Coverage is uneven: TypeScript supplies 68 PRs and Python 41. Together that is (68 + 41) ÷ 219 = 49.8%, a calculation from the published metadata. C has one PR and C++ two; an overall score gives much less evidence for those stacks.
Under the inspected repository classifier rubric, relevant, concrete findings in pre-existing or unchanged code can count, as can useful low-severity improvements. A true positive therefore does not necessarily mean a newly introduced, merge-blocking bug. For historical results, confirm the rubric version rather than assuming every row used identical scope rules.
Before comparing agents, decide whether your team wants only new defects or also maintenance suggestions. Specify which categories warrant an interruption, how much triage time is acceptable and which languages and change sizes must be represented.
Read the denominators before the rank
ReviewBench separates known-issue coverage and comment quality into different metrics. The scoring implementation makes the distinction explicit:
Metric | What it counts | How to use it |
|---|---|---|
Grounded precision | Known-valid matched findings ÷ all matched candidate findings. Unmatched comments are excluded. | Check quality among recognized findings; not noise across all comments. |
Grounded recall | Known valid reference findings covered ÷ all known valid reference findings. | Compare coverage of the same known issues. |
Augmented precision | Matched valid findings plus judge-approved unmatched findings ÷ all candidate findings. | Inspect comment quality across the full output, including new findings. |
Augmented recall | Known valid findings covered plus new approved findings ÷ known valid findings plus new approved findings. | Use as a per-agent diagnostic; each agent can expand its own denominator. |
Hypothetical example, not an agent result: assume one PR has 10 known valid findings. A reviewer emits 10 unique comments. Four match reference findings: three valid and one invalid. Of the six unmatched comments, the judge accepts one. Assume the matches cover distinct reference findings.
Grounded precision is 3 ÷ 4 = 75%, but augmented precision is 4 ÷ 10 = 40%. Grounded recall is 3 ÷ 10 = 30%; augmented recall is 4 ÷ 11 ≈ 36.4%. The gap shows why a seemingly precise reviewer may still produce substantial noise outside its matched comments.
For low-noise inline review, inspect augmented precision, actual comments per PR and representative false positives alongside grounded recall. For a security-focused role, inspect relevant category and severity misses; the score does not certify a repository as secure. Matching and novel-finding validity still depend on model judgments.
The methodology’s F-beta weighting lets beta below 1 favor precision and beta above 1 favor recall. Changing that setting re-ranks existing results; it does not change how an agent reviews code. Record the metric family, severity/category selection, aggregation, model configuration and snapshot date when sharing a comparison.
Separate adapter checks, private tuning and official results
A buyer can start by inspecting results and example findings on the leaderboard. Builders who want to evaluate their own reviewer should distinguish three stages:
Validate the adapter. Package a container that reads a PR snapshot and emits structured findings. The local runner needs Docker, Git and jq; start with one PR, then the 25-case subset. It validates output format but neither judges findings nor publishes a result.
Tune privately. Generate findings, then use the separate judging pipeline to score them. Keep the corpus revision, agent/model configuration, available context, budget and judge settings fixed when comparing changes. A local score with another judge is not an official leaderboard score.
Submit an official final. The onboarding guide requires a GitHub Container Registry image pinned by digest and registration through the submission portal. Maintainers approve onboarding. After testing, a final uses three fresh rounds over all 219 PRs; publication also requires maintainer review. The 25-case test result is not a leaderboard result.
Check environment compatibility before spending on a full run. The agent contract specifies linux/amd64, no GPU, declared HTTPS destinations and a default 15-minute limit per PR. Official runs use minimized repository snapshots, not unrestricted access to branches, history or live GitHub context. Local success with a fuller clone does not establish official-run compatibility.
The documented final implies 219 × 3 = 657 nominal PR-review executions, before retries or tuning. Under the current cost terms, tests and tuning are self-funded. ReviewBench pays final-submission judge inference, but the submitter still pays for the reviewer’s model access. The documentation does not establish an all-in dollar price.
Budget adapter work, failed attempts and human triage too. The reported review-duration metric measures the successful invocation for current container-based runs, excluding setup, earlier failures and subsequent judging. It is not end-to-end CI latency.
Make the adoption decision on your own repositories
The full reference labels are public, and the 25-case subset sits inside the 219-case corpus. Three fresh executions help expose run variability; they do not create unseen evaluation cases. The practical next step is a separately held-out, permissioned set of historical PR snapshots or shadow reviews of new changes.
For that pilot, keep answer labels and later fixes out of the reviewer’s context. Have maintainers adjudicate findings and misses against rules agreed before the run. A useful acceptance record should capture:
Useful unique findings and important missed issues, split by the categories and severities that matter to the repository.
False positives, duplicate comments and the time maintainers spend deciding what to act on.
Failures, end-to-end latency and total attempted-run cost—not only successful calls.
Adopt the configuration only for the role those results support. A low-noise advisory reviewer and a high-recall specialist may need different thresholds. Neither a strong benchmark row nor a good pilot replaces deterministic tests, domain checks or human acceptance. The distinction is also central to RohitAI’s guide to behavioral acceptance checks for agent-written migrations.
For teams maintaining an evaluation over time, the guide to versioned agent tasksets and reward audits explains why the task, scoring rules and execution environment must travel with a reported score.
Methodology: AI-assisted reporting and analysis based on GitHub’s announcement, public ReviewBench documentation and scoring code at revision ceb0794a, plus corpus metadata calculations. Availability and current repository files were rechecked on October 5, 2026. No reviewer, judging run or production pilot was executed; the numerical example is illustrative.
