AI productivity is often presented as a single number.
A tool makes developers 20% faster. An agent creates a 2x productivity gain. A team reports that AI saves several hours each week. Another study finds that experienced developers complete work more slowly with AI.
These results appear to contradict one another.
Often, they do not.
They may be measuring different developers, different tasks, different codebases, different stages of delivery, and different definitions of productivity. One number compresses all of those choices into a result that looks more universal than it is.
The question is not simply whether an AI productivity estimate is correct.
The question is: correct for which work?
One Number Can Be Accurate and Still Mislead
Consider three findings from AI productivity research.
Google studied more than 10,000 developers using a specific machine-learning code-completion system and reported a 6% reduction in coding iteration time. The measurement focused on the time between builds and tests within a bounded code-completion workflow.
METR’s early-2025 randomized study examined experienced open-source developers working on tasks in repositories they already knew. In that setting, participants took 19% longer when AI tools were allowed.
A later METR survey of 349 technical workers found a median self-reported increase of roughly 1.4x to 2x in the value of their work, while the median reported speed increase was 3x.
Six percent faster, 19% slower, and up to 3x faster cannot be placed on the same chart without context.
The Google study measured a specific part of the development loop. METR’s experiment measured task-completion time for experienced maintainers in familiar repositories. The survey measured participants’ perceptions of changes in speed and value across a broader range of technical work.
Each result answers a different question.
The mistake begins when one of them is presented as the productivity effect of AI for every engineering organization.
AI Productivity Depends on the Work
AI does not affect every software task equally.
It may generate boilerplate quickly, draft tests, explain an unfamiliar API, translate between frameworks, or prepare a first implementation of a bounded feature.
The same tool may struggle with a mature codebase whose important constraints are distributed across undocumented conventions, production history, architectural decisions, and the knowledge of a few experienced engineers.
Even within one repository, the productivity effect may differ across:
- feature development;
- bug investigation;
- test creation;
- refactoring;
- documentation;
- dependency migration;
- security-sensitive changes;
- architecture work;
- production incident response.
A team-wide average blends these categories together.
If AI creates strong leverage for repetitive migrations but adds review and correction work to architecture-sensitive changes, the average may still look positive. Yet the management decision should not be “use more AI everywhere.”
It should be “expand AI where leverage is durable and strengthen controls where it is not.”
AI Changes Which Tasks People Choose
The measurement problem becomes harder because AI does not only change how quickly people complete work.
It also changes which work they choose to do.
METR describes this through the idea of task substitution. When AI makes one category of work cheaper, people allocate more time toward that category. They may attempt tasks that previously seemed too expensive, too repetitive, or outside their expertise.
That creates at least three different productivity questions.
Uplift on old tasks
How much faster could the team complete the same basket of work it performed before AI?
This comparison holds the task mix constant. It is useful for understanding whether AI reduces the effort required for established work.
Uplift on new tasks
How much faster can the team complete the work it chooses after AI becomes available?
This task mix may include work the team would not previously have attempted. AI may make prototypes, migrations, internal tools, documentation, or broad test generation economically practical.
Uplift in value
How much more useful engineering value does the organization produce after people rearrange their work around AI?
This is the most important management question—and usually the hardest one to answer.
Under METR’s simplifying assumptions, uplift on old tasks is expected to be lower than uplift in value, while uplift on new tasks is expected to be higher. The exact result depends on how time is reallocated and how the organization defines value.
The distinction matters because completing a previously expensive task cheaply does not automatically mean that the task creates proportionally more business or engineering value.
AI can make new work possible. Measurement still has to ask whether that work was worth doing.
Selection Is Part of the Result
AI productivity studies also face a growing selection problem.
In METR’s later developer experiment, 30% to 50% of surveyed participants said they avoided submitting some tasks because they did not want to risk being assigned to complete them without AI.
This means the experiment was systematically less likely to include some of the work where developers expected AI to help most.
That does not prove the missing tasks would have produced the expected gains. Developers can overestimate AI’s effect. But it shows why a clean company-wide multiplier is difficult to establish.
Developers do not use AI randomly.
They choose it for tasks where they expect an advantage. They avoid it when explaining the problem to the tool would take too long, when the context is too sensitive, when the quality bar is difficult to verify, or when they already know the codebase better than the model can.
Observed AI productivity therefore contains two effects:
- how much AI helps on a given task;
- which tasks people choose to perform with AI.
A dashboard that records only AI-assisted output cannot separate them.
The Same Developer Can Have Several AI Leverage Profiles
The variation is not only between junior and senior developers or between different teams.
The same person can experience very different AI leverage across one working day.
A developer may use AI to generate test fixtures at 10:00, investigate a production race condition at 11:00, write documentation at 14:00, and review an architecture-sensitive pull request at 16:00.
AI may be highly effective for the fixtures, moderately useful for documentation, distracting during incident investigation, and valuable only as a second opinion during review.
Assigning that developer one AI productivity score hides the decisions the organization actually needs to make.
The useful unit of analysis is not simply the person or the tool.
It is the interaction between the person, task, repository, workflow stage, and resulting contribution.
Repository Context Changes the Outcome
A task that appears identical in a ticketing system may require very different engineering effort in different repositories.
Adding an endpoint to a new internal service is not the same as changing a payment flow in a mature production system. Updating a well-tested library is not the same as modifying legacy code with limited test coverage. Generating a new component is not the same as preserving compatibility across years of product behavior.
Repository context affects:
- how much information the AI can access;
- how much architectural knowledge remains implicit;
- how easy the output is to test;
- how expensive failure would be;
- how much review is required;
- how much rework appears later.
This is why AI adoption should not be evaluated independently from codebase maturity, test coverage, ownership, and maintainability.
A model can generate plausible code in both environments. The human effort required to trust that code may be completely different.
Speed at Creation Is Only One Measurement Point
AI productivity also changes depending on when it is measured.
At creation, the team may see:
- shorter time to first implementation;
- more code generated;
- more pull requests opened;
- faster completion of repetitive tasks.
During review and integration, the team may see:
- additional verification;
- more review iterations;
- architectural correction;
- duplicated implementations;
- test expansion;
- coordination with core maintainers.
After deployment, the team may see:
- follow-up fixes;
- rollback or replacement;
- reduced or increased incident risk;
- code that becomes a durable foundation;
- work that disappears after a short useful life.
A number measured at the prompt window cannot represent the entire lifecycle.
This is why AI productivity should be evaluated across delivery, rework, and durability—not only generation speed.
Segment First, Aggregate Second
Engineering leaders still need summary metrics. The answer is not to reject averages completely.
The answer is to earn the average by making the variation underneath it visible.
A useful analysis should segment AI impact across dimensions such as:
Work category
Does AI create different leverage for features, tests, refactoring, bug fixes, documentation, configuration, and security work?
Repository context
Does leverage change between new and mature repositories, well-tested and fragile systems, or internal and customer-critical services?
Developer familiarity
Does AI help most when a developer is exploring an unfamiliar area, or when an experienced maintainer can quickly judge the output?
Delivery stage
Does time saved during implementation remain saved after review, integration, testing, and deployment?
Time horizon
Does the contribution remain effective, or is the initial gain followed by rework, replacement, or removal?
Only after those differences are visible does a company-wide multiplier become useful.
The average should summarize the system. It should not conceal it.
A Hypothetical Example
Suppose an engineering organization reports an overall AI productivity multiplier of 1.7x.
That number sounds decisive.
The underlying pattern might look like this:
- test and fixture creation: 2.4x leverage with low rework;
- repetitive migrations: 1.8x leverage with stable results;
- feature implementation: 1.5x leverage with normal review effort;
- legacy bug fixes: 1.0x leverage with highly variable outcomes;
- architecture-sensitive work: 0.9x leverage after review and correction.
These figures are illustrative, not GitMe customer data.
The average is mathematically valid. But it does not tell leaders where to expand AI, where to improve context and tooling, or where human ownership should remain closest to the work.
The segmented view does.
It suggests increasing AI support for tests and migrations, studying what makes feature work successful, and applying stronger review or narrower use cases in legacy and architecture-sensitive areas.
One number reports the outcome.
The distribution explains what to do next.
A Better AI Productivity Scorecard
A stronger measurement system connects several layers.
AI participation
GitMe’s AI Effort Share shows where AI-assisted work is concentrated. It provides adoption context without treating usage as productivity.
AI leverage
AI Leverage expresses the relationship between modeled engineering output and modeled human effort as a multiplier.
This helps answer whether AI participation is associated with more productive capacity—not merely more generated activity.
Work composition
Work categories show what the organization produced: features, refactoring, tests, bug fixes, documentation, security, configuration, or other engineering work.
This is essential because the same aggregate multiplier can come from very different portfolios.
Rework
Review iterations and later corrective changes help show whether initial acceleration transferred effort downstream.
Rework should not automatically be classified as waste. Product learning, responsible refactoring, and changing requirements all create legitimate revision. The goal is to identify recurring patterns, not punish change.
Durability
Effort Survival Cohort adds the time dimension: how much earlier modeled engineering effort remains effective as the codebase evolves?
A contribution that supports later work has a different economic profile from one that is quickly rewritten or removed, even if both initially received the same productivity credit.
Together, these layers move AI measurement from one headline number toward an explainable system.
What Leaders Should Ask
Instead of asking only “How much faster are we with AI?”, leaders can ask:
- Which work categories produce the strongest AI Leverage?
- Where does high AI Effort Share fail to produce a meaningful multiplier?
- Which repositories require the most verification and correction?
- Does leverage change with developer familiarity?
- Are developers shifting toward work that AI makes cheap but that creates limited value?
- Does implementation speed survive review and integration?
- Which AI-assisted contributions require disproportionate rework?
- Does the modeled effort remain effective three, six, or twelve months later?
- Is the company-wide average being driven by a narrow set of tasks or a durable system-wide improvement?
These questions are harder than reading one percentage from a dashboard.
They are also much closer to the decisions engineering leaders actually have to make.
What GitMe Makes Visible
GitMe is designed to help engineering organizations interpret AI-assisted work with context.
By connecting modeled engineering effort, AI Effort Share, AI Leverage, work categories, historical comparison, rework, and Effort Survival Cohort, GitMe helps teams see both the summary and the distribution underneath it.
The purpose is not to assign every developer a simplistic productivity score. It is not to claim that one metric can fully represent engineering value.
The purpose is to distinguish:
- where AI participates;
- where it creates leverage;
- what type of work benefits;
- where human effort moves;
- what happens after delivery;
- how long the resulting effort remains effective.
A single AI productivity number can be useful.
It becomes useful only when the organization can explain what created it.
Related GitMe Reading
- AI Usage Is Not AI Leverage
- The Hidden Maintenance Tax of AI-Assisted Coding
- Engineering Productivity Needs a Half-Life
- AI-Generated Code Should Be Measured by How Long It Survives in Production
Sources
- METR. Task Substitution and Uplift. May 8, 2026.
- METR. We Are Changing Our Developer Productivity Experiment Design. February 24, 2026.
- METR. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. July 10, 2025.
- METR. Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity. May 11, 2026.
- DORA. Balancing AI Tensions: Moving from AI Adoption to Effective SDLC Use. March 10, 2026.
- Google Research. ML-Enhanced Code Completion Improves Developer Productivity. July 26, 2022.