The AI productivity gap is real, and it is far smaller than the marketing implies. That is not the problem. The problem is that most organisations cannot see where the missing difference actually goes.
The signal: 65% more AI, 8% more throughput
On 29 July, LeadDev published a number that deserved more argument than it got. Justin Reock, Deputy CTO at DX, reported that across more than 400 companies tracked from November 2024 to February 2026, AI tool usage rose 65% while median pull request throughput rose 7.76%.
His framing was generous, and I think it was correct: the benchmark is wrong, not the results. A 5–15% throughput gain across 500 engineers is real money. Nobody should apologise for it.
However, set that against LeadDev’s own Engineering Leadership Report 2026, published in June from a survey of 600 engineering leaders. AI for internal use had become the number one engineering priority, cited by 77% of respondents, up from 45% a year earlier. For the first time in the study’s three-year history, it outranked adding new features and capabilities at 71%.
So the largest reallocation of engineering priority in three years now points at an intervention that moves the headline number by roughly eight per cent. Eight per cent compounds, and the ceiling sits higher than most organisations have reached. Still, four things follow from that gap, and none of them appear on a throughput chart.
Truth one: the AI productivity gap begins as a measurement gap
Ask the same 600 leaders how they assess the impact of AI coding tools, and the most common instrument turns out to be employee feedback about the tools, at 63%.

Everything more rigorous sits at roughly a fifth. Development time per feature reaches 31%. Change failure and pull request reversion rates reach 22%. Weekly time saved per developer reaches 22%. Time spent reviewing AI-suggested code reaches 21%.
Read those together and the shape of the problem becomes clear. The single largest engineering priority of 2026 runs on sentiment. Leaders are not lazy. Sentiment is simply the only instrument they already had installed.
Sentiment is not evidence
I have spent a long time in delivery environments where you cannot make a change of this size because the team says it feels quicker. Someone asks what the controlled measure is, what the failure signal looks like, and what would make you stop. Those questions are not bureaucracy. They let you discover you were wrong while it is still cheap.
When the only instrument you own is a survey, the only failure mode you can detect is unhappiness.
James Stanier, CTO at Nordhealth, makes the same point in the report. “Self-reported productivity has never been the best signal,” he says. “It should be paired with other signals, like the traditional DORA metrics (are we actually going faster?) and token spend (who is using the tools the most?).”
This matters more than it looks. A change failure rate measures the system. Employee feedback measures the humans. Choose only the human instrument and you do not stop measuring the cost. You simply move where it registers. I made a related argument about spend in Token Spend Is Not a Cost to Cut.
Truth two: people absorb the AI productivity gap
Here is where it registers instead.

Hands-on technical responsibilities increased for 37% of engineering leaders, up from 25% in 2025. Among staff, principal and distinguished engineers, that figure reaches 52%, up from 30%. Among CTOs, 44% now do more hands-on technical work.
Meanwhile 45% work more hours than a year ago, up from 38% in 2025 and 35% in 2024. Scope grew for 63%. Neither the number of direct reports nor the number of teams fell to compensate. The role got wider, the day did not get longer, so the week did.
A third of managers, 32%, now consider a move back to individual contributor work, up from 24%. In addition, 41% say their team feels less motivated than twelve months ago, although that figure improved slightly from 44%.
The free-text answers deserve a board’s attention. Leaders describe rewarding creative work giving way to the maintenance and review of AI-generated code. One respondent describes an internal leaderboard ranked purely on token usage. Another calls the new shape of the job “a babysitting chore, rather than doing interesting work.”
In organisations that measure AI with sentiment alone, the burnout numbers are the productivity telemetry.
None of that reaches a throughput chart. Yet all of it describes the same phenomenon the chart fails to capture. Individuals absorb the difference between what AI promised and what it delivered, because individuals remain the only part of the system anyone instruments well. The burnout figures arrive late, in the wrong units, and at the wrong committee. I explored a similar illusion in AI May Make Work Feel Faster Without Making It Faster.
A note on the number everyone quotes
The most-shared statistic from the report says 54% of CTOs feel emotionally drained at least once a week, against 24% in 2025. That is a thirty-point swing in a year, and I think the direction holds.
Even so, the base deserves honesty. CTOs make up 6% of the 600 respondents, or roughly 36 people. The report publishes no per-question response counts, and 530 of the 600 finished every question. Hold a thirty-point movement on a sample that small loosely.
I raise this because it repeats the same failure one level up. We critique an evidence problem using evidence thinner than the argument needs. The sturdier claim is duller and harder to dismiss: working hours rose across every role in the survey, including a jump among advanced engineers from 28% to 53%. Quote that one instead.
Truth three: verification is real work, and nobody has costed it
Plenty of people have now written the piece explaining that AI moved the constraint from writing code to reviewing it. I wrote one myself in May: AI Coding Has Moved the Bottleneck From Creation to Verification. The more useful question is why, if the constraint moved so visibly, almost nobody can point to it in their own numbers.
The answer is that AI pushed effort into a part of the delivery system with no capacity line against it.
What the pull request data shows
LinearB’s 2026 Software Engineering Benchmarks, built from more than 8.1 million pull requests across 4,800 teams in 42 countries, found that agentic AI pull requests wait 5.3 times longer for reviewer pickup than unassisted ones. That is 1,055 minutes against 201. Acceptance rates for AI-generated pull requests run at 32.7%, against 84.4% for manual ones. Once a human finally starts, they review roughly twice as fast. The delay is queueing, not comprehension.
A study of 309 engineering leaders and practitioners by the platform Flux, reported by LeadDev in June, found that 35% of teams write code with AI but will not ship it, because they cannot assess the risk confidently. Only 3.6% said AI-generated issues never reach production.
CloudBees’ 2026 State of Code Abundance, based on research with more than 200 enterprise technology leaders, states the contradiction cleanly. Ninety-two per cent expressed confidence in the production readiness of AI-generated code. Eighty-one per cent reported an increase in production issues tied to it.
Kris Kang, chief product officer for agent systems at JetBrains, describes the mechanism: “We are effectively running code generation at machine speed, while the downstream verification and organizational processes are still dragging along at human speed.”
The consequence is mundane and expensive. Review, verification, security assessment and production support have all grown. In most organisations, none of them received a named owner, a capacity allocation or a service level. Goodwill and evenings fund them instead. A sentiment-only measurement regime cannot see that cost, whereas a change failure rate would surface it within a fortnight.
Truth four: the AI productivity gap is eating the junior pipeline
The clearest unpriced cost sits in the talent data.

Eighty-four per cent of respondents believe AI-powered tools will make it harder for junior developers to enter and grow in the profession. It ranks third among their concerns at 62%, behind code maintainability at 73% and output quality at 71%.
Now consider the response. Asked how AI most affects their approach to talent, leaders chose upskilling existing engineers at 49%, increasing productivity expectations at 18%, and rethinking hiring profiles at 12%. Supporting junior engineers in an AI-assisted world scored 4%.
To be fair to the data, this measures a belief about the future rather than the labour market itself. The report handles that distinction carefully elsewhere: only 8% say AI reduced headcount in 2025, against 21% who expect it to in 2026. Nevertheless, the direction of the talent response is not in doubt, and it carries an obvious consequence.
Verification capacity is seniority. Looking at a large AI-generated change and knowing which part is load-bearing is not a tool licence. It is accumulated exposure to systems failing in specific ways. Every organisation solving its verification problem by leaning harder on experienced people is spending a stock it has also decided not to replenish. That works for a few years. It does not work for ten.
Four practical changes I would make first
1. Instrument the cost side before you scale the agents
Two metrics, not twenty: change failure rate and rework rate. Both already have definitions. Both appear in the LinearB benchmark tables, so you can see where you sit. Both fail loudly. Twenty-two per cent of organisations track change failure rate today, and joining that group costs about a fortnight of plumbing.
2. Put verification on the plan as named capacity
Not “reviews happen”. Give it a percentage of team capacity, an owner, and a queue you can see. The LinearB pickup-time figure describes a queueing problem, and you solve queueing problems by allocating servers rather than by asking people to try harder. If review takes 20% of the work, it takes 20% of the plan.
3. Split the two questions your board keeps merging
“Is AI making us faster?” and “Is AI making us safer to ship?” have different answers and different instruments. Report one blended number and you get exactly what CloudBees found: 92% confidence sitting alongside 81% more production incidents.
4. Make the junior decision explicitly
If your honest position is that you will not hire and develop juniors for two years, say so. Then plan where senior verification capacity comes from in 2031. You can revisit a deliberate decision. You cannot revisit a drift, because nobody ever made it.
What the next two years look like
None of this argues for slowing AI adoption, and I would not make that case. The 7.76% is real, the ceiling sits higher, and the organisations that reach it will be the ones that fixed the conditions rather than bought more tokens. I made the related point about system constraints in Speeding Up the Builder Won’t Clear the Jam.
But there is a version of the next two years where the gain is genuine, the reported numbers stay green, and the forty-five per cent working longer weeks quietly fund the whole difference alongside the juniors nobody hired. That version does not announce itself. It surfaces three years later as a maintainability problem, with nobody left who remembers why the system took that shape.
Measuring the AI productivity gap properly is not an accounting exercise. It is the difference between finding that out now and finding it out then.
Sources used
- Justin Reock, “AI productivity gains are closer to 10% than 10x”, LeadDev, 29 July 2026. Analysis of 400+ companies, November 2024 to February 2026.
- LeadDev, The Engineering Leadership Report 2026, in partnership with Postman. Survey of 600 engineering leaders, 31 March to 16 April 2026; 530 completed every question.
- LinearB, 2026 Software Engineering Benchmarks Report. 8.1m+ pull requests, 4,800 teams, 42 countries.
- Chantal Kapani, “AI-generated code sparks production confidence crisis”, LeadDev, 30 June 2026. Flux study of 309 engineering leaders and practitioners.
- CloudBees, The 2026 State of Code Abundance Report, 19 May 2026. Research with 200+ enterprise technology leaders.
Continue the journey
One thought leads to another.
Scroll to explore
Enterprise Delivery
Minimum Idea State: A Practical Way to Test Risky Assumptions
Leadership & Culture
How to Build Products People Love: One Risky Bet
Enterprise Delivery
Stop Owning People. Start Owning Outcomes
Enterprise Delivery
Speeding Up the Builder Won’t Clear the Jam: Why Most Digital Transformations Fail to Deliver Impact
AI Governance
Generative AI Customer Insights: What AI Can and Can’t Do
The 1% Book Shelf
Difference Isn’t the Landmine. Justification Is.


Leave a Reply