Engineering Velocity Benchmark Report 2026: The State of AI in Production
Nine datasets on what changed between writing code and shipping it. Where the median team lands, where the benchmark sits, and which hires close the distance.
Throughput, year over year
CircleCI 2026 · n = 28,738,317 workflows · median team

Engineering output rose in 2026 while delivery did not. Feature-branch throughput is up 59% year over year, main-branch throughput for the median team fell 7%, and only 5% of custom enterprise AI pilots reach production.
Every number on this page has been published somewhere. They just haven’t been read together.
So we pulled the nine datasets engineering leaders keep citing at each other in Slack, put them on one screen, and looked for the shape they make. It isn’t subtle. Far more code gets written. Less of it lands in front of a customer. That gap is the whole report.
KORE1 has staffed engineering teams since 2005, and our engineering staffing desk hears the same story from the hiring side. We get far fewer calls that start with “we need three more backend developers.” We get a lot more that start with “our review queue is forty pull requests deep and nobody owns the pipeline.”

Output is up. Delivery is not.
CircleCI analyzed 28,738,317 workflows and found average throughput up 59% year over year, the largest jump since the report started running in 2019. Then it split the number by branch. Feature-branch activity for the median team grew 15%. Main-branch throughput fell 7%.
Code is getting written. It isn’t landing.
Read from a different angle, LinearB’s 8.1 million pull requests across 4,800 organizations in 42 countries tell the same story, because the teams leaning hardest on AI merged 98% more of them while watching review time climb 91%, which nets out near a 10% organizational gain instead of the doubling raw output implies. Ten percent. Not one hundred.
Main-branch success rates tell you where the difference went. They dropped to 70.8%, the lowest in more than five years, against a recommended benchmark of 90%.

The bottleneck moved somewhere expensive
Reviewing is the job now.
The average team runs a seven-day cycle time from first commit to production, and four of those days sit in review. That’s 57% of the clock spent deciding whether something is safe to merge rather than building it.
AI-assisted pull requests make it worse in a specific, measurable way. They’re bigger. At the 75th percentile they run 408 lines against 157 for unassisted work, agentic pull requests wait 5.3 times longer for a reviewer to pick them up, and acceptance splits hard once someone finally reads them, with bot-authored code landing 32.7% of the time against 84.4% for human-authored code.
Recovery slipped too. The median team now takes 72 minutes to get back to green after a failed main-branch run, up 13% year over year against a 60-minute benchmark. Mid-sized companies fare worst, approaching three hours, roughly four times the recovery time of both the smallest and the largest cohorts. Nobody planned for that.
Four numbers that define 2026
Where the median team sits, and where the benchmark sits
Six delivery measures with a published benchmark or a published comparison. Instrument one. Make it the first row.
Scroll the table sideways to see every column →
| Measure | 2026 median | Benchmark or comparison | Gap | Source |
|---|---|---|---|---|
| Main-branch success rate | 70.8% | 90% | −19.2 pts | CircleCI 2026 |
| Recovery to green after a failure | 72 min | 60 min | +12 min | CircleCI 2026 |
| Cycle time, first commit to production | ~7 days | <25 hrs (elite) | ~6× | LinearB 2026 |
| Share of cycle time spent in review | 57% | — | 4 of 7 days | LinearB 2026 |
| Pickup time, agentic AI vs. unassisted PRs | 1,055 min | 201 min | 5.3× | LinearB 2026 |
| Merge acceptance rate, bot vs. human authored | 32.7% | 84.4% | −51.7 pts | LinearB 2026 |
Cycle-time tiers, first commit to production
LinearB 2026 Software Engineering Benchmarks Report · n = 8.1M pull requests, 4,800 organizations, 42 countries
Adoption is not deployment, and deployment is not return
Read this table top to bottom. It’s the same funnel every AI program walks down, and almost nobody finishes it.
Scroll the table sideways to see every column →
| Stage | Share | Source and sample |
|---|---|---|
| Developers using or planning to use AI tools | 84% | Stack Overflow 2025 · n = 49,000+ |
| Technology professionals using AI at work | 90% | DORA 2025 · n ≈ 5,000 |
| Organizations using AI in at least one function | 88% | McKinsey, Nov 2025 · n = 1,993 |
| Organizations running AI agents in production | 52% | Google Cloud 2025 · n = 3,466 |
| Organizations attributing any EBIT impact to AI | 39% | McKinsey, Nov 2025 · n = 1,993 |
| Organizations reporting AI is fully scaled | 7% | McKinsey, Nov 2025 · n = 1,993 |
| Custom enterprise AI pilots that reach production | 5% | MIT Project NANDA 2025 |
The trust gap underneath it
Usage ran ahead of confidence, and it stayed there. Stack Overflow’s 2025 survey found 46% of developers actively distrust AI accuracy against 33% who trust it, with 3.1% saying they trust it highly. DORA’s read is consistent, with 30% reporting little or no trust in AI-generated code even though more than 80% say AI made them more productive.
Usage is not trust. Both of those are true at once, and the combination is what a backed-up queue looks like from the inside. A reviewer who doesn’t trust the diff in front of them reads every line of it, which is exactly the behavior you want and also exactly the behavior that turns a 98% jump in merged pull requests into a queue nobody can clear.
The most uncomfortable finding isn’t a survey at all. METR ran a randomized controlled trial with 16 experienced open-source developers across 246 real tasks on repositories they already knew well. Developers using AI took 19% longer to finish. Afterward, they estimated AI had made them 20% faster.
That’s a 39-point gap between what a delivery dashboard shows and what an engineer believes. We pulled that thread further in our breakdown of AI copilot adoption versus measured productivity. Surveys measure feelings. Delivery dashboards measure delivery. Worth remembering the next time someone quotes you a productivity number from the first kind.
What the benchmarks say about who to hire
Four gaps in the data. Four different people. Hiring more authors into a review-constrained team makes the queue longer, not shorter.
Senior engineers who review
When 57% of cycle time sits in review, the constraint is reading capacity. Screen for it directly.
Software engineer staffing →Platform and DevOps engineers
A 70.8% main-branch success rate is a CI/CD ownership problem before it’s a code problem.
DevOps engineer staffing →Site reliability engineers
Seventy-two minutes to green against a 60-minute benchmark is a rollback and observability gap.
SRE staffing →ML platform engineers
The 5% that ship have someone who owns evaluation, deployment, and monitoring.
ML platform engineer staffing →
What this looks like on the requisition side
Our engineering desk fills roles in 17 days on average, and 92% of the engineers we place are still there at twelve months. Those two numbers matter more together than apart. Speed alone is a trap. A fast bad hire on a review-constrained team makes the queue worse, because somebody still has to read their pull requests.
The mix has shifted hard since early 2025. Staff-level reviewers, platform engineers with real CI/CD ownership, and MLOps people who have actually shipped a model past evaluation are the three hardest searches we run across 30-plus U.S. metros. Not because candidates don’t exist. Because the screen changed and most job descriptions still test for authorship.
Two companion pieces go deeper on the AI side of this. The AI pilot-to-production checklist covers what to verify before you commit headcount, and our build versus buy breakdown covers whether you should be hiring for it at all. On the delivery side, the Clarity Stack covers three reasons velocity stalls that have nothing to do with tooling.
Every source, with its sample
No KORE1 data is blended into the industry figures. Where we cite our own numbers, they’re labeled as ours and they come from our placement records.
- 01
- 02
- 03
- 04
- 05
- 06
- 07
- 08
- 09
McKinsey’s State of AI survey (published November 5, 2025, n = 1,993 respondents across 105 countries) is cited for the adoption, EBIT and scaling figures. KORE1 figures throughout, the 17-day average time-to-hire and the 92% twelve-month retention rate, come from our own placement data across eight verticals.
Common Questions
Did AI actually make engineering teams faster in 2026?
Individually yes, organizationally barely. Teams with heavy AI adoption merged 98% more pull requests, but review time rose 91% and net organizational productivity gain landed near 10%, per LinearB’s 2026 benchmark set of 8.1 million pull requests. Throughput and delivery came apart.
What counts as a good cycle time for an engineering team?
Under 25 hours from first commit to production is elite, 25 to 72 hours is good, 73 to 161 hours is fair, and anything over 161 hours needs focus. Those tiers come from LinearB’s 2026 report. The average team runs about seven days, so most organizations sit in the bottom two bands and don’t know it. Instrument cycle time before you argue about it.
Why do AI-generated pull requests sit in the queue longer?
Mostly because they’re bigger and less trusted. Agentic AI pull requests wait 5.3 times longer for reviewer pickup, at 1,055 minutes against 201, and AI-assisted PRs run 408 lines at the 75th percentile against 157 for unassisted work. Reviewers also accept bot-authored code 32.7% of the time versus 84.4% for human-authored code, so the expected payoff for opening one is lower. Smaller PRs fix more of this than better prompts do.
How many AI pilots actually make it to production?
About 5% of custom enterprise AI tools reach production, according to MIT’s Project NANDA research covering 300-plus disclosed initiatives. Meanwhile 52% of organizations report at least one AI agent running in production per Google Cloud’s 2025 ROI study, and 74% of those say they saw return within the first year. The split is build versus buy. General-purpose tools cross the line, bespoke builds mostly don’t.
What does DORA say about AI and delivery stability?
DORA’s 2025 report found AI adoption positively related to throughput and negatively related to software delivery stability. Roughly 90% of technology professionals now use AI at work and more than 80% report a productivity gain, while 30% say they have little or no trust in the code it produces. DORA frames AI as an amplifier that magnifies whatever practices a team already has.
Which roles actually fix a review bottleneck?
Reviewers, platform owners, and reliability engineers, in that order. Adding authors to a team where 57% of cycle time is review makes the queue longer. We screen for review throughput and rollback experience directly rather than inferring it from years on a résumé, which is a different interview than the one most teams are running.
How long does it take to hire an engineer who can close these gaps?
KORE1’s average time-to-hire across engineering roles is 17 days, and 92% of the engineers we place are still in seat at twelve months. Platform and MLOps searches trend longer, usually three to four weeks, because the qualifying screen is narrower. Contract, contract-to-hire, and direct hire all run out of the same bench across 30-plus U.S. metros.
Hire for the part that’s actually broken
Tell us where delivery is stalling, whether that’s review capacity, pipeline ownership, recovery, or a model stuck one step short of production. We’ll come back with engineers who have fixed that exact failure.
17-day average time-to-hire · 92% twelve-month retention · 30+ U.S. metros
