Up next
Home
/
Library
/
AI in software development
/
Is our AI coding investment paying off?

Is our AI coding investment paying off?

Photo of Jamie Birss
|

Summary

  • An AI coding investment is paying off when merged, working code grows relative to spend, not when adoption or token spend rises.
  • Per LinearB CTO Yishai Beeri, human-led PRs yield in the high 80s to low 90s, while fully agentic flows fall to around 30%.
  • Cost per PR blends human and AI spend against merged output, giving finance one number to trend quarter over quarter.
  • Read yield rate and cost per PR together; either one alone produces a false read.

Is our AI coding investment paying off?

Your AI coding investment is paying off if: 

  1. Your code merged
  2. Working code is growing relative to what you spend 

Better adoption or increased token spend are not appropriate ROI measurements. Those two numbers move independently of each other constantly, and conflating them is the most common measurement mistake engineering leaders make with AI right now.

We dive deeper into that distinction below by looking into PR yield rate and why that's the cleanest signal available to measure your AI ROI. 

Why adoption and spend don't answer the question

Adoption tells you a tool is installed. Spend tells you money moved. Neither tells you whether the work that came out the other end was worth having.

Yishai Beeri, CTO at LinearB, has spent recent months in LinearB's own pull request data trying to separate the signal from the noise, and his framing is blunt about where the confusion comes from.

"We're seeing a very big divide in the kind of productivity gains that different teams and different organizations and even different developers get."

— Yishai Beeri, CTO at LinearB, on Dev Interrupted, The playbook to close your team's AI productivity gap

The gap he's describing shows up on the same adoption dashboard. Two teams can report identical AI usage and land in completely different places on delivered output, because usage was never the thing standing between spend and value. Writing code faster only helps if the code survives review, doesn't need rework, and actually ships. A team that codes twice as fast but reworks a third of it hasn't doubled anything.

Nik Sudan, engineering operations lead at Kraken, reaches the same conclusion from a different direction, and states it as bluntly as a metric gets stated.

"One word, four letters, data. If you're not measuring AI effectiveness, adoption cost, and relating to that as well, the quantitative engineering data like code throughput, DORA, then it's all worthless, 100% worthless."

— Nik Sudan, engineering operations lead at Kraken, on Dev Interrupted, How Kraken finds hidden bottlenecks across thousands of engineers

What PR yield rate actually shows

PR yield rate is the percentage of opened pull requests that actually get merged, and it is the cleanest signal available for whether AI-assisted work is converting into shipped code. LinearB's benchmark data, drawn from millions of pull requests, shows this number falling in a specific, informative pattern as human ownership drops out of the loop.

Human-authored work yields high, typically in the high 80s to low 90s. A human using AI but staying in charge of the pull request sees only a small dip, two or three points, attributable to larger diffs and harder reviews. Fully autonomous agentic flows, where an agent pulls a ticket and pushes a PR with no human driving it to merge, collapse to around 30 percent.

"The agent creates three PRs, only one will actually make it to the code base. And I think the obvious attribution here is to lack of ownership."

— Yishai Beeri, CTO at LinearB, on Dev Interrupted, The playbook to close your team's AI productivity gap

That's the diagnostic value of the metric. A low yield rate doesn't mean the model is bad. It means nobody is chasing approvals, answering review comments, or pushing the change to merge, because detaching the human from the PR detaches the accountability that gets code shipped. Fixing ownership, not swapping models, is what moves this number.

What cost per PR actually shows

Cost per PR blends human and AI spend into one figure, divided by pull requests merged, and it is the number Beeri points to when a leader wants a single line the CFO can track quarter over quarter.

"If you're looking at your cost per PR... that represents the leverage you're getting from this new technology, this new tool, this new kind of like electricity running through your factory."

— Yishai Beeri, CTO at LinearB, on Dev Interrupted, The playbook to close your team's AI productivity gap

Sudan's version of the same idea works at the level of an individual contributor rather than the whole org, and it's a sharper way to catch the failure mode this metric is built to expose.

"You take two engineers who have the same output and the same impact. One spends thousands of dollars and the other hundreds. Who's the most effective engineer here?"

— Nik Sudan, engineering operations lead at Kraken, on Dev Interrupted, How Kraken finds hidden bottlenecks across thousands of engineers

The engineer burning thousands of dollars for the same merged, impactful output as the one spending hundreds is simply more expensive, not more effective. Cost per PR is the number that makes that visible, where raw output or raw spend alone would hide it.

The two metrics only work together

Yield rate and cost per PR check different failure modes, and reading one without the other produces a false read in both directions. High yield with a climbing cost per PR means the work is landing but getting more expensive to land, usually a sign of rework eating the gains. Low cost per PR with a falling yield rate means spend looks efficient on paper while an increasing share of that work never actually merges.

Watching both together is what turns a feeling into a number you can defend to finance.

Frequently asked questions

Does high AI adoption mean the investment is working?

No. Adoption measures whether the tool is installed and used, not whether the resulting work merges or holds up. Teams with identical adoption numbers can land at completely different points on delivered output.

Why does agentic PR yield collapse so much more than AI-assisted human coding?

Ownership. A human working with AI still chases their own approvals and pushes the change to merge. A fully autonomous agent that opens a PR and moves on leaves nobody responsible for getting it across the line, so unreviewed work piles up and gets abandoned.

What's a healthy cost-per-PR trend to look for?

Falling or flat as throughput rises. A rising cost per PR alongside rising output usually means rework is quietly consuming the gains AI appeared to produce.

Is token spend itself a useful number?

On its own, no. It tells you money moved, not whether that spend produced merged, working code. Pair it with yield rate and cost per PR before drawing any conclusion from it.

Photo of Jamie Birss

Jamie Birss

Jamie is a product marketer at LinearB with a soft spot for great technology and the humans who make it. Outside of work, you'll usually find him out running a trail, or shipping his next puzzle game.

Connect with

Your next read