# When AI Reviews AI, Who Owns the Merge?

> AI can now write a pull request, review another agent’s work, and help revise the result within minutes. But automated agreement does not create accountability for what reaches production.

- Published: 2026-09-29
- Canonical: https://www.gitdotme.com/blog-when-ai-reviews-ai-who-owns-the-merge

A coding agent opens a pull request. An AI reviewer scans the diff, identifies potential problems, and leaves comments. The authoring agent responds, modifies the implementation, and updates the tests. Continuous integration turns green.

From the outside, the workflow looks complete.

The code was written. The code was reviewed. The feedback was addressed. The checks passed.

But one question remains:

Who owns the decision to merge?

This is becoming a practical engineering-management problem. AI is no longer present only on the creation side of software development. It is increasingly operating on both sides of the pull request: one system produces the change, while another evaluates it.

That can make delivery faster. It can also create a closed loop in which automated agreement is mistaken for independent judgment.

The output of an AI review is evidence. It is not accountability.

## The Closed Loop Already Exists

AI-to-AI code review is not a hypothetical future workflow.

A 2026 study by Niruthiha Selvanayagam and Taher A. Ghaleb linked AI-attributed pull requests with AI-attributed review events on GitHub. Their dataset contained 248,641 unique AI-attributed pull requests that received at least one AI-attributed review.

Of those pull requests:

- 208,145 received a review from the same product family;
- 45,269 received a review from a different product;
- 4,773 appeared in both groups.

The authors also found that cross-product AI-to-AI review remained a minority of identified agent-authored activity, but its volume increased by more than two orders of magnitude between the first and third quarters of 2025.

The study does not prove that these reviews were correct, independent, or responsible for the eventual merge decision. Attribution methods, repository selection, reviewer composition, and timestamp availability all limit what can be concluded.

But the scale establishes something important: AI-generated work is already being evaluated by other AI systems in real development workflows.

The engineering question is no longer whether AI can review AI.

It is what that review should be allowed to mean.

## A Second Model Is Not Automatically an Independent Reviewer

Teams often treat multiple automated opinions as if they create independence.

An authoring agent proposes a change. A reviewing agent criticizes it. The first agent responds. Agreement between them appears to increase confidence.

Sometimes it should.

A second system may notice a missing test, an unsafe assumption, an unhandled edge case, or a mismatch between the pull request description and the actual diff. Automated review can provide valuable coverage at a speed human teams cannot match manually.

But using two systems does not automatically create independent judgment.

The agents may share:

- similar training data;
- similar assumptions about common coding patterns;
- incomplete repository context;
- the same misleading pull request description;
- the same missing business requirement;
- the same blind spot around production behavior;
- incentives optimized for producing a plausible answer rather than accepting operational responsibility.

Even cross-product review does not guarantee independence. Two different tools can still rely on the same incomplete evidence.

True independence has at least three dimensions:

1. **Different reasoning paths:** the reviewer should not simply reproduce the author’s assumptions.
2. **Different evidence:** tests, runtime behavior, policy checks, historical context, and repository constraints should challenge the proposed change.
3. **Separate authority:** someone must have the responsibility to accept the remaining risk.

AI can strengthen the first two dimensions.

The third remains an organizational decision.

## Review Activity Is Not a Review Outcome

AI review dashboards can create another measurement trap: counting comments as proof of quality.

A review agent may produce twenty observations. The authoring agent may resolve all twenty. The pull request may show a clean conversation with no unresolved threads.

That still does not establish that the important risks were examined.

A large-scale study of AI code-review actions analyzed more than 22,000 review comments across 178 repositories. The researchers found that effectiveness varied considerably. Concise comments, comments containing code snippets, manually triggered reviews, and feedback attached to specific hunks were more likely to lead to code changes.

That finding is useful, but it also highlights an important distinction.

A comment leading to a change shows that the review affected the diff. It does not prove that the change was correct, that the highest-risk issue was identified, or that the resulting implementation will behave well in production.

Comment count measures activity.

Resolution count measures workflow completion.

Neither, by itself, measures accountable engineering judgment.

## Green Checks Are Necessary—but They Do Not Own the Decision

Continuous integration provides stronger evidence than review volume. Tests, builds, linters, security scans, and policy checks can challenge a change with repeatable rules.

They should be central to AI-assisted development.

But a green pipeline answers only the questions encoded into that pipeline.

It cannot confirm a requirement that was never written. It cannot protect an architectural constraint it was not designed to test. It cannot recognize that a technically correct feature should not exist. It cannot decide whether the residual risk is acceptable to the business.

Research on 33,596 agent-authored pull requests illustrates this broader integration problem. Across the studied repositories, 71.48% of the pull requests were merged. Documentation, CI, and build-related tasks had the highest overall merge rates, while performance and bug-fix work had lower acceptance.

Pull requests that were not merged tended to contain larger changes, touch more files, fail more CI checks, and require more review revision. In a qualitative analysis of 600 rejected pull requests, reviewer abandonment was the most frequent pattern. CI or test failure appeared in 17% of the examined rejections, alongside duplicate submissions, unwanted features, incorrect implementations, and misalignment with reviewer instructions.

These results come from open-source repositories and should not be treated as a universal benchmark for every engineering organization.

Their larger lesson is still relevant: integration failure is not purely a code-generation problem. It is also a coordination, intent, review, and ownership problem.

## Merge Success Is Not the Same as Production Success

A merged pull request has passed an organizational boundary.

It has not completed its economic life.

The change may still:

- trigger a rollback;
- require an urgent correction;
- create repeated review or maintenance work;
- duplicate an existing capability;
- increase operational complexity;
- be rewritten shortly after release;
- survive for months and become a durable part of the product.

This is why the person responsible for the merge cannot be defined only as the person who clicked the button.

Ownership includes accepting responsibility for:

- the intended outcome;
- the evidence used to approve the change;
- the risks the evidence did not eliminate;
- the production monitoring plan;
- the rollback path;
- the downstream consequences.

AI can help assemble this evidence. It cannot hold organizational accountability for the outcome.

## Human Review Should Mean Ownership, Not Manual Repetition

Keeping a human accountable does not require a senior engineer to manually reread every generated line with no automated assistance.

That approach would not scale, and it would discard much of the value AI can provide.

Human control should instead be designed around explicit ownership.

The accountable reviewer should be able to answer:

- What user or business outcome is this change intended to produce?
- What parts of the implementation were AI-assisted?
- Which assumptions were tested automatically?
- Which risks remain outside the automated checks?
- Was the reviewer given enough repository and domain context?
- Who will respond if the change behaves differently in production?
- How can the change be reverted safely?

The goal is not to force humans to duplicate the AI reviewer.

The goal is to ensure that someone understands why the available evidence is sufficient for this particular change.

## Not Every Pull Request Needs the Same Gate

A documentation correction should not require the same approval process as an authentication change, billing calculation, database migration, or production access policy.

Teams need risk-adjusted review paths.

### Low-risk, bounded changes

Examples include documentation, formatting, isolated test additions, or well-defined maintenance tasks.

These changes may rely heavily on automated authoring, automated review, and CI, with a lightweight human confirmation or an explicit policy for automatic integration.

### Medium-risk product changes

Feature logic, shared components, integrations, and behavior changes need stronger evidence.

AI review can perform the first pass, but an accountable engineer should confirm intent, test coverage, architectural fit, and the production verification plan.

### High-risk changes

Security, authorization, payments, customer data, infrastructure, destructive migrations, and compliance-sensitive behavior require named human ownership and domain-specific review.

For these changes, a second AI opinion should supplement—not replace—independent technical and organizational authority.

The objective is not maximum friction.

It is proportional accountability.

## Every Pull Request Needs an Evidence Chain

As AI participates in more stages of development, the pull request should evolve from an activity record into an evidence record.

A useful evidence chain includes:

1. **Intent evidence**\
   The issue, requirement, acceptance criteria, and reason the change should exist.

2. **Provenance evidence**\
   Which parts were produced or modified with AI assistance, which tools participated, and what context they received.

3. **Validation evidence**\
   Tests, CI results, security checks, static analysis, runtime verification, and relevant manual inspection.

4. **Review evidence**\
   Which findings were raised, which were accepted or rejected, and which decisions required human judgment.

5. **Ownership evidence**\
   The named person or role accepting the remaining risk and authorizing the production boundary.

6. **Production evidence**\
   Deployment verification, incidents, rollback events, follow-up fixes, rework, and the durability of the resulting contribution.

An AI reviewer can help create and organize this evidence chain.

It should not be allowed to erase the distinction between evidence and authority.

## Measure the Chain, Not the Number of Approvals

Adding more automated reviewers can create the appearance of stronger governance while making the workflow noisier.

Three AI approvals are not necessarily more valuable than one well-configured reviewer, a meaningful test suite, and a clearly accountable engineer.

Teams should measure what happens across the full chain:

- author provenance;
- reviewer provenance;
- human review effort;
- review iterations and corrective work;
- CI and policy failures;
- time to responsible approval;
- post-merge fixes and reversions;
- production incidents;
- contribution durability.

These measures should be used to improve the engineering system, not to create a simplistic score for individual developers.

The purpose is to identify where automation reduces effort without reducing control—and where apparent speed is transferring risk to another phase.

## What GitMe Makes Visible

GitMe is designed to help engineering organizations evaluate AI-assisted work beyond activity counts.

AI Effort Share shows where AI participation is concentrated. AI Leverage examines the relationship between modeled engineering output and modeled human effort. Rework patterns and historical comparisons help reveal whether initial acceleration creates additional stabilization work. Effort Survival Cohort adds the time dimension by showing how modeled engineering effort continues to remain active in the evolving codebase.

These signals do not decide whether a pull request should be merged.

They help leaders evaluate what happened after the decision:

- Did AI participation reduce the human effort needed to deliver the work?
- Did review and corrective effort rise elsewhere?
- Did the approved contribution survive?
- Did the organization create durable capacity or only accelerate activity?
- Did the review system produce reliable evidence, or merely more comments?

Accountability remains human.

Measurement makes the consequences visible.

## Accountability Begins Where Automated Agreement Ends

AI reviewing AI is not inherently a failure of engineering governance.

Used well, it can improve coverage, shorten feedback loops, detect routine problems earlier, and allow human reviewers to focus on intent, architecture, risk, and production impact.

The danger begins when automated agreement is treated as if it removes the need for an owner.

A pull request can be generated by AI.

It can be reviewed by AI.

It can be revised by AI.

It can pass every automated check.

But when that change crosses into production, the organization still needs to know who accepted the evidence, who accepted the remaining risk, and who will own the result.

The merge button is not only a workflow action.

It is an accountability boundary.

## Related GitMe Reading

- [Why One AI Productivity Number Is Never Enough](https://www.gitdotme.com/blog-why-one-ai-productivity-number-is-never-enough)
- [AI Usage Is Not AI Leverage](https://www.gitdotme.com/blog-ai-usage-is-not-ai-leverage)
- [The Hidden Maintenance Tax of AI-Assisted Coding](https://www.gitdotme.com/blog-hidden-maintenance-tax-ai-assisted-coding)
- [AI-Generated Code Should Be Measured by How Long It Survives in Production](https://www.gitdotme.com/blog-ai-generated-code-survives-production)

## Sources

- Selvanayagam, N., & Ghaleb, T. A. [AI-to-AI Code Reviews of GitHub Pull Requests](https://arxiv.org/abs/2608.21311). arXiv, 2026.
- Kamalı, H. Ö., Tuna, E., Haratian, V., & Tüzün, E. [Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review](https://arxiv.org/abs/2605.17548). arXiv, 2026.
- Ehsani, R., et al. [Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub](https://arxiv.org/abs/2601.15195). MSR 2026.
- Sun, K., et al. [Does AI Code Review Lead to Code Changes? A Case Study of GitHub Actions](https://arxiv.org/abs/2508.18771). arXiv, 2025.
- Nachuma, C., & Zibran, M. [When AI Teammates Meet Code Review: Collaboration Signals Shaping the Integration of Agent-Authored Pull Requests](https://arxiv.org/abs/2602.19441). arXiv, 2026.

## Make AI Accountability Visible

Connect AI participation, modeled human effort, review, rework, and contribution durability with GitMe.

[Get Started](https://panel.gitdotme.com/signup)

Analyze repositories, surface real contribution signals, and measure developer performance with context instead of vanity metrics.
