Back to Insights
10 Oct 20262 min readAI Agents

Agents write the code. Review is now the bottleneck.

Coding agents can open pull requests faster than any team can read them. The work did not disappear. It moved to review.

LinearB's benchmark report across more than 8.1 million pull requests puts a number on it: AI-written pull requests wait 4.6x longer before anyone picks them up, though they are reviewed about 2x faster once someone does (LinearB). The same report finds an acceptance rate of 32.7% for AI-generated pull requests against 84.4% for manual ones.

Why "faster" can be slower

It is tempting to measure agents by how much code they produce. A randomized study by METR is a useful warning: 16 experienced open-source developers working on repositories they knew were 19% slower with AI tools, while they believed they were about 20% faster (METR). It is a small study on specific tools, and METR lists its own caveats, but the gap between how fast it felt and how fast it was is the point. If review time is not measured, a team can feel faster while shipping slower.

What teams that run agents at scale do

Keep a person on every merge. Stripe reports more than 1,300 agent pull requests a week, all reviewed by a human (InfoQ). It also caps CI retries at two rounds, so a failing agent stops instead of looping (ByteByteGo).

Make waiting visible. Meta tracks time in review as a metric and tested nudges for reviews that sat too long; across more than 30,000 engineers, the nudges cut time in review by 6.8% (FSE 2022).

Use machine help where it is cheap to check. At Google, 7.5% of reviewer comments were resolved by an ML-suggested edit (Google Research). At Meta, AI-suggested patches for review comments were applied 19.7% of the time when they were actionable, and showing them to reviewers slowed reviewers down, so Meta showed them to authors instead (arXiv). The lesson is not that AI review fails. It is that who sees the suggestion matters.

A starting checklist

  1. Measure time to first review for agent pull requests separately from human ones. If agent PRs wait much longer, more agents will only lengthen the queue.
  2. Cap work in progress per engineer, not per agent. The limit is how much a person can carefully read.
  3. Cap retries. An agent that cannot get CI green in two rounds should stop and ask.
  4. Gate merges and production steps on a human, and make that step required, not advisory.
  5. Send machine suggestions to the author first, so reviewers see a cleaner diff instead of more noise.

The takeaway

The useful question is no longer how much code agents can write. It is how much a team can review well. Measure that, plan for it, and add agents only as fast as review can keep up.

AI agentscode reviewengineering metrics

Found this useful?