Thought Leadership

AI Usage Is Not AI Leverage

By GitMe Team •

Adoption metrics show how often teams use AI. They do not show whether AI produces more durable engineering output for the same human effort.

AI adoption has become one of the easiest numbers in engineering to measure—and one of the easiest to misinterpret.

Leaders can count licensed seats, active users, prompts, accepted suggestions, generated lines, or the percentage of developers using an AI coding tool. These numbers answer a legitimate question: is the organization using AI?

They do not answer the more important one: is AI allowing the organization to produce more durable engineering output for the same human effort?

That distinction is the difference between AI usage and AI leverage.

Usage measures participation. Leverage measures the change in productive capacity. A team can have high usage and low leverage, or modest usage and high leverage. Treating those states as equivalent can turn an adoption dashboard into a misleading account of productivity and return on investment.

Adoption Has Become the Baseline

The software industry is no longer asking whether developers will use AI.

DORA's 2025 research found that 90% of technology professionals were using AI at work and more than 80% believed it had increased their productivity. At the same time, DORA reported that higher AI adoption was associated with both higher software delivery throughput and higher delivery instability.

That combination matters. More output can coexist with a less stable delivery system.

DORA's later analysis describes AI as an amplifier. It can accelerate strong engineering systems, but it can also amplify weak feedback loops, fragmented knowledge, poor internal platforms, and inconsistent quality controls. Time saved during initial creation may reappear as time spent auditing, verifying, integrating, or correcting the result.

An adoption rate cannot show that transfer of work. It records that AI entered the process, not what happened to the process afterward.

Usage Metrics Answer the Wrong Question

Common AI dashboards tend to emphasize what is easy to count:

  • how many developers activated a tool;
  • how frequently they used it;
  • how many suggestions they accepted;
  • how many lines were attributed to AI;
  • how much code or how many pull requests the team produced.

These signals can help evaluate rollout and engagement. They are useful for license management, enablement, and identifying where AI is or is not being adopted.

But they do not establish productivity.

An accepted suggestion may save ten minutes, create thirty minutes of review work, or prevent several hours of debugging. A generated implementation may be production-ready, or it may duplicate an existing capability that the model could not see. A developer may use AI constantly for low-value tasks while another uses it selectively for work with much higher leverage.

The usage event looks similar in each case. The engineering outcome does not.

This is why adoption should be treated as an input metric. It describes exposure to the tool. It does not measure the value created through the tool.

Perceived Speed Is Not the Same as Measured Gain

Research on AI-assisted development appears contradictory when different measurements are treated as if they were interchangeable.

In an early-2025 randomized controlled trial, METR studied experienced open-source developers working on issues in repositories they already knew. With the AI tools available at the time, participants took 19% longer to complete the assigned work—even though they expected AI to make them faster.

That result should not be generalized to every developer, tool, repository, or task. The study examined a particular group of experienced maintainers, a particular period in rapidly changing AI tooling, and work performed in familiar codebases.

METR's follow-up work also found evidence that later tools may provide productivity benefits. But the researchers encountered a measurement problem: as AI became more useful and more widely adopted, developers became less willing to perform randomly assigned no-AI work. They also became more selective about which tasks they submitted to the experiment, often withholding tasks they particularly wanted to complete with AI.

The implication is not that one study is right and the other is wrong. It is that AI's effect depends on the developer, task, tool, codebase, and measurement method.

Self-reported speed, controlled task completion time, repository output, and long-term maintainability each observe a different part of the system. None should be presented as the entire return.

High AI Users May Already Be High Performers

Another measurement trap appears when teams compare heavy AI users with everyone else.

GitClear's 2026 research reports that heavy users of AI coding tools produced four to ten times more output than non-users. Taken alone, that difference could be interpreted as an enormous AI productivity multiplier.

But much of the gap existed before those developers adopted AI. When heavy users were compared with their own earlier performance, the reported gain was closer to 25%.

The 25% gain may still be valuable. The important point is methodological: a cross-sectional difference between two groups is not the same as the improvement caused by a tool.

High-performing developers may adopt new tools earlier. They may receive more suitable tasks, have stronger repository knowledge, work in better-supported teams, or simply be trusted with a larger volume of changes. If these factors are not separated from AI use, existing performance differences can be misreported as AI leverage.

Counting who uses AI most tells you where adoption is concentrated. It does not tell you how much additional capacity AI created.

AI Effort Share and AI Leverage Measure Different Things

GitMe separates AI participation from the productivity multiplier associated with it.

AI Effort Share indicates where AI-assisted work is concentrated within modeled engineering effort. It helps teams understand how much of the work includes AI participation and how that participation is distributed across people, repositories, or periods.

AI Leverage asks a different question: how much modeled engineering output is being produced relative to the modeled human effort required to produce it?

That relationship is expressed as a multiplier. A value such as 1.5x is not a statement that 50% of the code was written by AI. It indicates modeled productive leverage: the relationship between delivered engineering effort and the human effort estimated to remain necessary.

The distinction can be summarized simply:

  • AI Effort Share measures adoption within the work.
  • AI Leverage measures the productivity multiplier associated with that work.

Neither should be interpreted alone.

High AI Effort Share with limited leverage may indicate that the tool is widely used but is not materially reducing human effort. Lower AI Effort Share with stronger leverage may indicate selective use on tasks where AI creates more value.

The goal is not to maximize AI participation. It is to find where AI produces useful, repeatable, and durable leverage.

A Simple Comparison

Consider two hypothetical teams.

Team A reports that 65% of its modeled engineering effort includes AI participation. It produces more code and opens more pull requests than before. But reviews require more iterations, senior engineers spend more time correcting integration problems, and many recent changes are rewritten within weeks.

Team B reports 35% AI Effort Share. Its developers use AI more selectively: for test generation, unfamiliar APIs, repetitive migrations, and bounded implementation tasks. The team delivers a smaller increase in raw code volume, but comparable work requires less modeled human effort, rework remains controlled, and the resulting contributions continue to support the codebase over time.

Team A has higher AI usage. Team B may have higher AI leverage.

An adoption dashboard would reward Team A. A productivity system should investigate the full outcome before deciding which team is creating more value.

Leverage Must Survive Review, Rework, and Time

Even a genuine short-term multiplier is incomplete if the resulting work is expensive to stabilize or disappears quickly.

AI leverage should therefore be interpreted alongside at least two downstream dimensions.

The first is rework. Teams should examine whether AI-assisted contributions require more revisions, follow-up fixes, reversions, or replacement than comparable work. Rework is not automatically waste; requirements change and responsible engineering includes iteration. The useful question is whether the pattern changes systematically as AI participation rises.

The second is effort durability. GitMe's Effort Survival Cohort examines how modeled engineering effort remains effective in the evolving codebase over time. It helps distinguish work that continues to support later development from work that is quickly rewritten, removed, or replaced.

Together, these measures prevent a common accounting error:

  1. record the speed gained during code creation;
  2. ignore the effort spent making that code production-ready;
  3. ignore how long the contribution remains useful.

AI leverage is strongest when the output requires less human effort to deliver, does not create disproportionate downstream work, and continues to survive in the codebase.

A Better AI Productivity Scorecard

No single metric can explain AI's effect on an engineering organization. A useful scorecard connects four layers.

1. Participation

Where is AI being used, and what share of modeled engineering effort includes AI participation?

This is the role of AI Effort Share. It establishes the adoption context without treating adoption as the outcome.

2. Leverage

How much modeled output is produced relative to modeled human effort?

This is the role of AI Leverage. It tests whether AI participation is associated with a meaningful productivity multiplier.

3. Stabilization

What happens before and after delivery?

Review iterations, corrective work, reversions, follow-up changes, and the distribution of that work across the team reveal whether local speed is transferring cost to another phase or another engineer.

4. Durability

How much of the modeled engineering effort remains active over time?

Effort Survival Cohort adds the time dimension that adoption and throughput metrics lack. It helps leaders evaluate whether apparent leverage becomes lasting engineering value.

These layers should be examined together, with relevant historical baselines and context. They are not a formula for ranking developers. They are a way to understand whether the engineering system is becoming more productive.

What Leaders Should Ask

The shift from usage to leverage changes the management conversation.

Instead of asking only how many developers use AI, leaders can ask:

  • Where does AI participation produce the strongest leverage?
  • Which task types show little or no measurable gain?
  • Does higher AI Effort Share reduce modeled human effort, or merely increase output volume?
  • Are review and corrective workloads rising with AI-assisted output?
  • Is downstream work becoming concentrated among senior engineers?
  • Do AI-assisted contributions survive as long as comparable work?
  • Is the organization buying durable capacity or temporarily accelerating code production?

These questions make room for a more realistic result. AI can be highly valuable without being equally valuable everywhere. A tool can improve one category of work and slow another. A team can become faster at implementation while becoming less efficient at integration or maintenance.

The purpose of measurement is not to force one universal conclusion. It is to discover where leverage is real.

What GitMe Makes Visible

GitMe is designed to help engineering organizations move beyond activity counts and adoption dashboards.

By connecting modeled engineering effort, AI participation, AI Leverage, work categories, rework patterns, historical comparisons, and Effort Survival Cohort, GitMe helps teams evaluate whether AI-assisted development creates durable productive capacity.

The system is not intended to replace engineering judgment, act as a simplistic employee score, or claim that every contribution can be reduced to one perfect number.

Its role is to make distinctions that conventional dashboards blur:

  • adoption versus leverage;
  • output volume versus modeled human effort;
  • initial delivery versus downstream stabilization;
  • work produced today versus effort that continues to survive.

Those distinctions are essential for evaluating AI return on investment.

Adoption Is the Beginning, Not the Result

AI usage is becoming normal. That makes usage itself less informative as a measure of competitive advantage.

The meaningful question is no longer whether developers have access to AI or how often they invoke it. The question is whether AI changes the relationship between human effort and durable engineering output.

Measure participation—but do not confuse it with productivity.

Measure the multiplier—but do not ignore the work required to stabilize it.

Measure what ships—but also measure what survives.

AI usage tells you that a tool is present. AI leverage tells you whether the engineering system became more capable.

Sources

Measure leverage, not just adoption

Connect AI participation, modeled human effort, rework, and effort durability with GitMe.

Get Started