AI coding tools can create a strange management experience.
Developers generate code faster. Commit activity rises. More pull requests arrive. AI usage spreads across the organization. Yet releases do not accelerate at the same rate, delivery queues remain full, and the expected business impact is difficult to find.
None of these observations necessarily contradicts the others.
They measure different parts of the system.
AI may accelerate the activity closest to the developer while leaving review, integration, testing, deployment, adoption, and maintenance constrained by the same human and organizational bottlenecks as before. A faster coding step does not automatically create a faster engineering organization.
The central question is therefore no longer whether AI can help someone write code faster.
It is how much of that local acceleration survives the journey from code generation to released, used, and durable software.
The Productivity Gain Shrinks on Its Way to Production
A September 2026 revision of the NBER working paper Writing Code vs. Shipping Code studied more than 500,000 GitHub developers together with AI usage telemetry.
The researchers examined three generations of AI coding tools:
- Autocomplete
- Interactive coding agents
- Autonomous coding agents
Their matched event-study estimates found cumulative increases in commit activity of approximately 30%, 180%, and 240%, respectively.
The largest upstream gain did not propagate through the rest of the production system at the same rate. For autonomous agents, the estimated 240% increase in commit activity fell to 80% at the project level and 30% for actual releases.
| Production layer | Estimated cumulative effect for autonomous agents |
|---|---|
| Commit activity | +240% |
| Projects | +80% |
| Releases | +30% |
This is not evidence that the remaining productivity disappeared completely. A 30% release increase would still be material. It is evidence that the unit being measured changes the conclusion.
If leaders look only at commits, they may see a transformation.
If they look at releases, they may see a meaningful but much smaller improvement.
If they continue toward product adoption, reliability, and long-term maintenance, the effect may change again.
The NBER analysis is observational rather than a randomized controlled trial, so its percentages should not be treated as universal causal estimates for every organization. Its production hierarchy is nevertheless valuable because it shows where apparent gains can attenuate.
The paper estimates an elasticity of substitution of 0.23 between AI and human effort, consistent with strong complementarities in the production chain.
The closer measurement gets to customer value, the more of the wider engineering system it must include.
Individual Speedups Are Real, but They Are Not Universal
Other research reinforces the need to separate local task performance from organizational performance.
Randomized field experiments at Microsoft, Accenture, and an anonymous Fortune 100 company included 4,867 developers. When the three experiments were pooled, access to an AI coding assistant was associated with a 26.08% increase in completed tasks, with a standard error of 10.3%. Less experienced developers showed higher adoption and larger gains.
That is credible evidence that AI can improve task-level productivity in real organizations.
But another randomized study produced a very different result in a different setting.
METR studied 16 experienced open-source developers completing 246 real tasks in mature repositories they knew well. Participants had worked on their respective repositories for an average of five years. They expected AI to reduce task completion time by 24% before the study and still believed it had made them 20% faster afterward.
Measured completion time showed the opposite: access to the early-2025 AI tools increased completion time by 19%.
METR explicitly cautioned against generalizing this finding to every developer, tool, or task. The environment involved experienced maintainers working in large, familiar repositories, and AI capabilities continue to change quickly.
Taken together, the studies do not support a simple claim that AI always accelerates or always slows software work.
They support a more useful conclusion:
AI productivity depends on the developer, task, codebase, workflow, tool, and outcome being measured.
That is why an organization cannot infer its own result from a benchmark, a vendor claim, or an industry average.
It needs internal evidence.
The Bottleneck Moves Instead of Disappearing
Software delivery is a connected system.
A change may pass through requirements, architecture, implementation, review, testing, integration, release approval, deployment, adoption, operations, and maintenance. Accelerating one stage increases the amount of work presented to the next one.
If code generation becomes dramatically cheaper while review capacity remains fixed, the review queue grows.
If review accelerates but testing remains constrained, work waits there instead.
If more software reaches production but users do not adopt it, the organization has increased output without increasing value.
DORA’s 2025 research describes AI as an amplifier: it magnifies the strengths and weaknesses already present in an organization. The largest returns come from improving the surrounding organizational system, not merely purchasing better tools.
GitLab’s 2026 AI Accountability survey offers a similar signal. Among 1,528 developers and technology buyers, 79% agreed that individual developer productivity had improved with AI while the overall software delivery process had not accelerated at the same pace. Eighty-five percent agreed that the bottleneck had shifted from writing code to reviewing and validating it.
Those figures describe respondents’ perceptions rather than a causal experiment. They still capture an operational pattern engineering leaders should investigate inside their own teams.
When authoring capacity expands, downstream capacity becomes more important.
Activity Is Not the Same as Flow
AI adoption dashboards often begin with what is easiest to count:
- Active AI users
- Prompts or tokens
- Suggestion acceptance
- AI-attributed lines of code
- Commits
- Pull requests
These signals are useful. They show adoption and activity.
They do not reveal whether work moved through the system more efficiently.
A team can produce more pull requests while review time increases. It can merge more code while rework rises. It can ship more releases while customers use none of the new functionality. It can report a short-term output gain while the introduced code is replaced or removed a few months later.
The measurement system should follow work farther downstream.
| Measurement layer | Question | Useful signals |
|---|---|---|
| Adoption | Are developers using AI? | Active users, usage frequency, AI-attributed work, tool cost |
| Local leverage | Is AI increasing useful output relative to effort? | Accepted work, task completion, AI Leverage |
| Delivery flow | Does work move through engineering faster? | Review wait, cycle time, rework, release frequency |
| Production outcome | Does the change work after release? | Incidents, reversions, defects, customer adoption |
| Durable value | Does the contribution remain useful over time? | Code survival, replacement, continued modification, Effort Survival Cohort |
| Economic result | Is the return greater than the cost? | Modeled value, AI spend, operating cost, realized product impact |
No single row can replace the others.
The purpose is not to create the largest possible dashboard. It is to prevent an upstream activity metric from being presented as a downstream business outcome.
AI Usage Is Only the Starting Signal
An AI usage rate answers a narrow but legitimate question: how much of the observed engineering activity involved AI?
It does not answer whether AI improved performance.
High usage can coexist with strong leverage, weak leverage, or even negative leverage. Developers may use AI effectively for repetitive implementation work but lose time when applying it to unfamiliar architecture, complex debugging, or mature systems with undocumented constraints.
This is why GitMe separates AI usage from AI Leverage.
AI Leverage expresses the productivity multiplier associated with AI-assisted work. It asks whether AI use is translating into greater engineering output rather than treating adoption as the result itself.
That distinction matters for management decisions.
If usage is low and leverage is high, the opportunity may be broader enablement.
If usage is high and leverage is low, the organization may need better task selection, training, context, review practices, or tool governance.
If both are high, the next question is whether downstream delivery and durability confirm the same benefit.
Usage starts the investigation. It should not end it.
Shipping Is Still Not the Finish Line
A release is closer to business value than a commit, but it is not the final measurement point.
Software continues to generate evidence after deployment.
It may remain active, receive follow-on development, enable other work, and become part of the product’s durable foundation. It may also be reverted, rewritten, replaced, abandoned, or preserved only because nobody feels safe changing it.
A productivity measurement that ends on merge day cannot distinguish between those outcomes.
GitMe’s Effort Survival Cohort follows the share of earlier modeled engineering effort that remains active as the codebase evolves. It is not employee retention. It is a durability view of engineering contribution.
This adds time to the productivity question.
Instead of asking only, “How much did the team produce this month?” leaders can also ask:
- How much of that effort remained active after 30, 90, or 180 days?
- Did AI-assisted work survive differently from other work?
- Which teams converted AI usage into lasting contribution?
- Where did apparent speed become later replacement or rework?
A short-term acceleration that disappears in the following months should not receive the same interpretation as work that continues creating value.
Measure the Conversion, Not Just the Increase
Engineering leaders do not need to choose between optimism and skepticism about AI.
They need to measure conversion across the system.
A practical review should follow five questions:
- Did AI increase useful engineering output?
Compare usage with AI Leverage rather than assuming that adoption equals improvement. - Did the additional output move through delivery?
Examine review queues, cycle time, rework, testing, and release frequency. - Did the released work behave well in production?
Track defects, incidents, reversions, and operational burden. - Did the contribution remain valuable?
Follow code survival and Effort Survival Cohort over time. - Did the organization create more value than it spent?
Compare AI tools, supervision, review, correction, and maintenance costs with the value of the resulting work.
This creates a conversion chain:
AI usage → useful leverage → delivery flow → production outcome → durable value
Every transition can amplify or absorb the gain created upstream.
The goal is not to force every metric upward. The goal is to identify where the organization’s expected AI return stops propagating.
A Better AI Productivity Review
A credible AI productivity program should begin before the next tool rollout.
Establish a baseline for comparable teams and repositories. Segment results by task type, developer experience, codebase maturity, and AI workflow. Pair leading indicators such as usage with lagging indicators such as releases, production outcomes, and survival. Review the pattern over time rather than declaring success from a short pilot window.
Most importantly, treat conflicting evidence as useful.
More commits with slower reviews identify a capacity mismatch.
Higher AI usage with unchanged leverage identifies an adoption-quality problem.
Faster releases with weaker survival identify a durability problem.
Strong leverage with stable quality and survival identifies a repeatable advantage.
The organization becomes faster only when the gain moves through the system instead of accumulating at its first bottleneck.
AI can make developers faster.
The company gets faster when that speed becomes shipped, used, and durable value.
Related GitMe Reading
- AI Usage Is Not AI Leverage
- Why One AI Productivity Number Is Never Enough
- AI-Generated Code Should Be Measured by How Long It Survives in Production
- The Hidden Maintenance Tax of AI-Assisted Coding
Sources
- Demirer, Musolff, and Yang — Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools
- Cui et al. — The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers
- METR — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- DORA — State of AI-Assisted Software Development 2025
- GitLab — 2026 AI Accountability Research
Turn AI Activity Into Durable Engineering Value
AI adoption is visible quickly. Its real engineering return takes longer to prove.
GitMe helps engineering leaders distinguish AI usage from AI Leverage, benchmark engineering performance, and follow how modeled effort survives as the codebase changes.
Start with GitMe