AI Economics

The Hardest Question in AI ROI: Compared to What?

By GitMe Team • August 28, 2026

The Economist recently described corporate AI measurement as moving from "tokenmaxxing" toward "something more normal." That captures an important shift in enterprise AI: usage is no longer enough. Leaders now want returns.

But the moment a company asks for ROI, a harder question appears:

Compared to what?

Most AI ROI conversations still start in the wrong place.

Leaders ask whether developers are using AI, whether accepted code is increasing, whether cycle time improved, whether license utilization is high, and whether people feel faster. Those are useful questions. They are not the hardest question.

AI ROI is not the value of AI in isolation. It is the difference between one future and another. What happened with AI compared with what would have happened without it, or compared with a different way of spending the same money, attention, and engineering capacity?

If that comparison is weak, the ROI number will be weak too. It may still look precise in a spreadsheet, but it will not be decision-grade.

Adoption is not the counterfactual

AI usage dashboards are getting more sophisticated. GitHub's Copilot usage metrics, for example, expose adoption, engagement, acceptance, lines of code, and pull request lifecycle signals. That helps leaders understand whether a rollout is active and where AI-assisted work is entering the delivery flow.

But usage does not answer the counterfactual question.

A team can have high active usage and still create no net ROI if review load increases, rework rises, production defects grow, or the saved authoring time is absorbed by meetings, backlog churn, and coordination delays. A team can also have modest usage and strong ROI if AI removes a specific bottleneck in a high-value workflow.

The dashboard can show that AI was used. It cannot, by itself, prove that AI changed the outcome.

The Economist's useful baseline point

The Economist's setup is that usage must give way to outcome and return measurement. Our extension is that a return is only interpretable relative to a credible alternative or baseline.

Accessible summaries of The Economist's article highlight the same attribution problem: avoided hiring can be easier to connect to financial return than incremental revenue, which may require baselines or A/B testing. For engineering leaders, the pattern is familiar. Some benefits are visible because a team can handle the same workload with fewer planned hires or less outside help. Others are harder because delivery, quality, customer behavior, and revenue all move at once.

This is the bridge from the Economist's setup to the question this article focuses on. Once leaders stop asking "how much AI did we use?" and start asking "what return did we get?", they immediately need to define the comparison.

The evidence already points in both directions

This is why AI ROI is difficult. The evidence is not one-note.

Microsoft Research's controlled GitHub Copilot study, published in 2023, found that developers with Copilot completed a JavaScript HTTP server task 55.8% faster than the control group. That is a real productivity signal in a defined task setting.

METR's 2025 randomized controlled trial found a different result in a different setting: experienced open-source developers working on real issues in repositories they knew well took 19% longer when AI tools were allowed. METR also cautioned against over-generalizing the result and later changed its experiment design because newer studies faced selection and measurement problems.

DORA's generative AI research adds another layer. It reports that AI can improve individual developer well-being and perceived productivity, while also associating higher AI adoption with lower delivery throughput and delivery stability in the studied data.

McKinsey's August 25, 2026 State of AI survey shows the same enterprise tension. Eight in ten respondents said AI improved their individual productivity, while the share reporting enterprise-level EBIT impact from AI remained essentially unchanged from the prior year at 37%. McKinsey also found that about one in five respondents said AI-related operating costs, including token costs, constrained usage.

None of this proves that AI works or does not work in every engineering organization. It proves something more useful: ROI depends on the comparison, the workflow, the baseline, and the measurement window.

The wrong baselines make AI look better than it is

AI ROI often gets overstated because leaders compare AI-assisted work against an unrealistically weak baseline.

"With AI, this task took two hours" is not an ROI claim unless the company knows how long the same task would have taken without AI, at similar quality, with similar review standards, and with the same developer context.

"We generated 40% more code" is not an ROI claim unless the company knows whether that code was needed, reviewed, retained, and connected to business value.

"Developers feel faster" is not an ROI claim unless the organization can see whether the delivery system became faster after review, testing, release, and rework.

The weakest baseline is a memory of how slow things used to feel. The strongest baseline is a measured comparison against a plausible alternative.

The right baseline depends on the decision

There is no single universal AI ROI baseline. The correct comparison depends on what decision the leader is trying to make.

If the decision is whether to renew AI coding seats, the comparison is not "AI versus nothing." It is "this AI rollout versus a disciplined non-AI improvement plan with the same budget." That might include better CI, faster review, clearer requirements, reduced meetings, or developer enablement.

If the decision is whether to expand agentic coding, the comparison is not only "agent PRs versus human PRs." It is "agent-assisted delivery versus the best available human-plus-tooling workflow for the same class of work."

If the decision is whether to reduce hiring, the comparison must include future capacity, knowledge transfer, review pressure, apprenticeship, incident risk, and maintenance burden. A short-term cost reduction can look positive while long-term engineering capability declines.

If the decision is whether to build software internally with AI instead of buying it, the comparison must include ownership cost. McKinsey's 2026 survey found that 32% of respondents said their organizations decided against buying at least one software product or feature because they could build it internally with agentic coding tools. That may be rational. It also makes the counterfactual harder, because "buy" and "build" have different cost curves, risk surfaces, and maintenance obligations.

A practical counterfactual stack

Engineering leaders can make AI ROI more honest by building a counterfactual stack instead of a single headline number.

Start with the historical baseline. What did similar work cost before AI, measured in cycle time, review time, defect rate, rework, incident impact, and retained contribution?

Add a matched-work baseline. Compare similar repositories, teams, work types, and risk levels. A greenfield UI task should not be compared with a production incident fix in a legacy service.

Add a quality-adjusted baseline. AI-assisted work should be measured after review, tests, rework, production behavior, and later retention. First-pass speed is not enough.

Add an opportunity-cost baseline. What else could the organization have done with the same money and senior attention?

Add a time-window baseline. Accessible summaries of The Economist's article also point to a J-curve pattern: AI adoption may create early implementation cost or disruption before measurable gains appear. Some AI investments have an initial productivity dip, especially while teams learn tools, define policies, and reshape workflows. DORA's AI ROI materials similarly frame ROI as a rollout and budget conversation, not a one-day usage snapshot.

Finally, add an ownership baseline. Who maintains the generated system, explains the code, handles risk, and pays the future change cost?

The point is not to make measurement bureaucratic. The point is to stop pretending that AI ROI can be inferred from activity alone.

What to measure after choosing the comparison

Once the baseline is explicit, the measurement model becomes clearer.

For engineering AI, useful ROI signals include:

  • AI Effort Share: where AI meaningfully influenced the work, separated by team, repository, and work type.
  • Real Effort Value: how much durable engineering contribution was created after complexity, context, and quality are considered.
  • Review and verification cost: how much senior attention was needed to make the work safe.
  • Rework ratio: how often AI-assisted changes required correction after review, merge, or production release.
  • Contribution retention: how much of the shipped work survived later maintenance and refactoring.
  • Delivery impact: whether throughput, stability, lead time, customer value, or roadmap delivery improved.
  • Total AI cost: licenses, tokens, infrastructure, security work, training, governance, and operational overhead.

These signals do not replace financial ROI. They make the financial ROI credible.

Not all of these signals need to come from a single platform. Financial, operational, delivery, and engineering data may need to be combined to build the full ROI picture.

Where GitMe fits

GitMe provides engineering evidence that can support a more credible counterfactual analysis.

Historical comparison helps leaders compare engineering patterns across periods, while AI Effort Share shows where AI-assisted work appears in those periods. Real Effort Value helps separate meaningful contribution from raw activity. Rework helps reveal whether AI-assisted work is associated with less downstream correction or with more cleanup later in the delivery cycle. Contribution Retention helps answer whether shipped work remains useful after the next release, refactor, or maintenance cycle.

That matters because the question is not "Did developers use AI?" It is "What changed compared with the best realistic alternative?"

This connects directly with GitMe's views on why AI usage metrics still do not measure ROI, the real cost of AI-generated code, and the build-versus-buy-versus-own decision.

The better AI ROI question

The hardest question in AI ROI is not whether AI can help. The evidence shows that it can, in the right context.

The harder question is whether AI produced more durable value than the alternative use of the same money, time, talent, and attention.

Compared to no AI?

Compared to a better process?

Compared to better tests, faster reviews, clearer product decisions, or a smaller backlog?

Compared to buying software instead of building it?

Compared to hiring, training, or retaining stronger engineers?

AI ROI becomes serious only when leaders can answer "compared to what?" with evidence. Until then, the organization is measuring activity, not return.

Sources

Make AI ROI measurable against the right baseline.

Use GitMe to connect AI Effort Share, Real Effort Value, rework, contribution retention, and durable engineering outcomes.

Get Started