Up next
Home
/
Library
/
Engineering metrics
/
Why is our cycle time climbing?

Why is our cycle time climbing?

Photo of Jamie Birss
|

Summary

  • At Kraken, PR size and complexity explained only 5-10% of a review slowdown; 70-80% came from time zones and reviewer availability.
  • Review time is the stage of cycle time most likely to hide a people-and-process problem behind what looks like a tooling metric.
  • AI-assisted PRs do lower merge rates by two to three points, mostly due to size, but that's rarely the largest factor.
  • Diagnose in two layers: platform-level trends to locate the problem, then Git-provider detail to confirm the mechanism.

Why is our cycle time climbing?

Tracking delivery metrics helps us understand how things move. When delivery slows down, cycle time is often a good first place to start. Cycle time climbs for reasons that rarely match the obvious suspect. The common assumption is larger, more complex pull requests, often blamed on AI-generated code. 

Kraken's own investigation, detailed below, found that explains only a small fraction of it. The larger cause is almost always upstream of the code itself: who's available to review it, and when. That's the framework this page walks through, built on one team's investigation of their own numbers. 

The conventional explanation, and why it's usually wrong

The default story engineering leaders tell themselves is simple: AI is generating bigger, messier pull requests, reviewers are taking longer to get through them, and that's why cycle time is climbing. It's a plausible story. Kraken's own data says it's the wrong one more often than not.

Nik Sudan, engineering operations lead at Kraken, started with exactly that assumption.

"First of all, we assumed that big, complex merge requests were what was slowing the review time down. That's what most people say."

— Nik Sudan, engineering operations lead at Kraken, on Dev Interrupted, How Kraken finds hidden bottlenecks across thousands of engineers

His team tested it against real data instead of trusting the hunch, joining high-level LinearB metrics with granular signals pulled from their Git provider: reviewer activity, comment volume, files touched. The result reversed the assumption entirely.

"Size and complexity accounted for maybe 5 to 10% of the slowdown. 70 to 80% came from time zone differences... engineers in one region opening a merge request at a time that didn't line up with the code owners for the areas."

— Nik Sudan, engineering operations lead at Kraken, on Dev Interrupted, How Kraken finds hidden bottlenecks across thousands of engineers

The work wasn't stuck because it was hard to review. It was stuck because the person who could review it wasn't online yet.

Review time is the part of cycle time actually worth investigating

Deploy time and build time respond to infrastructure investment: faster pipelines, fewer dependencies, more runners. Review time responds to a different lever entirely, because Sudan frames it as a people-and-process problem rather than a tooling one. That's exactly why it's harder to diagnose, and why the surface-level guess tends to stop the investigation before it starts.

"Review time, in my view, is arguably the most important part of cycle time, and it is the current bottleneck for us and probably the bottleneck for many companies right now."

— Nik Sudan, engineering operations lead at Kraken, on Dev Interrupted, How Kraken finds hidden bottlenecks across thousands of engineers

If cycle time is climbing and deploy and build times are flat, review time is where to look first. It's the stage most likely to be hiding a people problem behind what looks like a tooling metric.

AI-assisted work does inflate this problem, just not the way the common assumption has it

There's a real mechanism connecting AI to cycle time, and per Beeri's data, PR size drives it more than PR complexity does. Yishai Beeri, CTO at LinearB, has watched this in LinearB's own benchmark data: even when a human stays fully in charge of an AI-assisted pull request, the merge rate ticks down slightly, and size is the specific reason why.

"The merge rates are gonna be only slightly lower, so two or three percentage points lower. And we can attribute that to the difficulty in reviewing AI codes, the larger PRs that typically you can, you typically get from AI."

— Yishai Beeri, CTO at LinearB, on Dev Interrupted, The playbook to close your team's AI productivity gap

That's a real effect, and worth fixing on its own terms. But per Kraken's own data, it's a two-to-three-point effect sitting on top of a seventy-to-eighty-point effect from reviewer availability. Chasing PR size first is optimizing the smaller variable while the larger one keeps climbing.

How to find your own version of this

The method that worked at Kraken generalizes past their specific finding. Start with the high-level trend: identify where cycle time is climbing using platform-level data like LinearB's, which tells you the shape of the problem without drowning you in row-level detail. Only then join that trend against granular Git-provider signals, reviewer timestamps, comment counts, code-owner assignments, to find the actual mechanism.

"The two-layer approach works really well: LinearB high, then your repository provider low, and then AI is the part that moves fluidly between each other and enriches it."

— Nik Sudan, engineering operations lead at Kraken, on Dev Interrupted, How Kraken finds hidden bottlenecks across thousands of engineers

Dumping every granular detail into the analysis from the start defeats the purpose. The high-level layer exists to point you at where to dig; the low-level layer exists to confirm or kill the hypothesis once you have one.

When this doesn't apply

If deploy time or build time are the ones climbing rather than review time, this framework isn't the fix. Infrastructure bottlenecks respond to infrastructure investment, not to reviewer scheduling. And if your team is genuinely small and colocated in one time zone, the reviewer-availability mechanism described here largely doesn't apply; look at PR size and review depth instead, since those are more likely to be the real driver in that specific setup.

Frequently asked questions

Is AI-generated code actually making cycle time worse?

It contributes, but modestly. In LinearB's benchmark data, AI-assisted PRs with a human still in charge see only a two-to-three-point dip in merge rate, attributable to size and review difficulty. That's real, but it's rarely the largest factor.

What should we look at first if cycle time is climbing?

Break it into deploy, build, and review time separately before doing anything else. Review time is the stage most likely to be hiding a people-and-process problem behind what looks like a tooling metric.

Why did Kraken's initial hypothesis turn out to be wrong?

It was a reasonable guess that matched conventional wisdom, and they tested it against real data instead of acting on it. Size and complexity explained a small share of the slowdown; time zone and reviewer availability explained most of it.

Photo of Jamie Birss

Jamie Birss

Jamie is a product marketer at LinearB with a soft spot for great technology and the humans who make it. Outside of work, you'll usually find him out running a trail, or shipping his next puzzle game.

Connect with

Your next read