# Your Coding Agent May Need Less Context, Not More

By GitMe Team • October 2, 2026

> Repository instructions can guide coding agents, but more context can also raise cost and encode stale assumptions. The goal is minimum sufficient context.

- Published: 2026-10-02
- Canonical: https://www.gitdotme.com/blog-coding-agent-may-need-less-context
- Author: GitMe Team

Every coding-agent failure can look like a context problem.

The agent edited the wrong file, so the team adds an architecture explanation.

It skipped an important test, so the command goes into `AGENTS.md`.

It misunderstood an environment boundary, so the deployment procedure is copied into the instructions.

It repeated an earlier mistake, so another warning is added.

Each addition is reasonable on its own. Over time, however, the instruction file becomes a second repository: part onboarding guide, part operations manual, part style guide, part incident history, and part list of rules nobody is confident enough to remove.

The underlying assumption is simple:

More context should produce better work.

For coding agents, that assumption is not always true.

## Repository Context Has Become Infrastructure

Repository-level instruction files are moving from an experimental practice to a normal part of software development.

The open `AGENTS.md` format describes itself as a README for coding agents and reports adoption across more than 60,000 open-source projects. GitHub also supports repository-wide, path-specific, and agent instruction files for Copilot workflows.

This adoption makes sense.

A coding agent entering an unfamiliar repository needs to know things it may not reliably discover from source code alone:

- which commands represent the real validation path;
- which directories are generated and should not be edited;
- where environment boundaries exist;
- which security constraints are non-negotiable;
- which actions require human approval;
- and what “done” means for that repository.

These instructions can prevent expensive mistakes.

But repository context is not free.

It consumes tokens, shapes search behavior, creates additional obligations, and can preserve assumptions long after the codebase has changed.

Once an agent is instructed to trust a rule, stale documentation becomes more than a documentation problem. It becomes executable organizational memory.

## The Evidence Is More Complicated Than “More Is Better”

Recent research does not support a universal answer.

Gloaguen and colleagues evaluated repository-level context files across multiple coding agents and language models. In the current version of their study, context files did not significantly improve task success overall, while increasing inference cost by more than 20% on average.

The agents did not simply ignore the instructions. They generally followed them.

They explored more files, performed broader testing, and spent more effort satisfying the additional requirements. Developer-committed context files outperformed LLM-generated ones by 7% on average, but repository overviews were not helpful and context files did not significantly improve performance overall.

The researchers recommend omitting automatically generated context files for now and keeping human-written files focused on non-standard practices or requirements not already documented in the README. Any claim that these files improve performance should be rigorously evaluated before deployment.

Another study reached a different operational result.

Lulla and colleagues examined 124 pull requests across ten repositories, comparing agent runs with and without `AGENTS.md`. They reported a 28.64% lower median runtime and 16.58% lower output-token consumption when the instruction file was present, while task-completion behavior remained comparable.

These findings are not mutually exclusive.

A concise file containing accurate commands and non-obvious constraints can reduce search.

A broad file containing generic explanations, duplicated guidance, obsolete decisions, and task-irrelevant rules can increase it.

The useful question is therefore not:

“Do context files work?”

It is:

“Which context helps which tasks, at what cost, and with what downstream result?”

## Context Bloat Is Already Measurable

The risks are not only theoretical.

A 2026 study of 100 popular open-source repositories containing `AGENTS.md` or `CLAUDE.md` files identified recurring configuration smells.

The researchers found:

- lint leakage in 62% of the analyzed files;
- context bloat in 42%;
- skill leakage in 35%;
- and frequent combinations of context bloat, skill leakage, and conflicting instructions.

Lint leakage occurs when instruction files repeat formatting and style rules that automated tooling could enforce more reliably.

Skill leakage occurs when detailed task procedures are placed in global instructions even though they apply only to a narrow workflow.

Context bloat occurs when the agent receives substantially more information than the task requires.

Each smell makes the instruction file appear more comprehensive. None guarantees better engineering output.

A long file can create confidence because it documents many cases. That confidence is dangerous if the agent is being anchored to outdated or irrelevant information.

## More Context Creates More Obligations

Humans do not treat every sentence in a large internal document as equally important.

Coding agents can behave differently.

A sentence placed in an instruction file may be interpreted as a requirement to satisfy, even when it is only background information. A historical workaround can become a permanent constraint. A recommendation can become a mandatory step. Two individually sensible rules can create a conflict the agent must resolve.

This produces several hidden costs.

### Wider exploration

Extra architectural detail can encourage the agent to inspect components unrelated to the requested change.

### Additional validation

A global list of commands may cause every task to run every check, even when only a small subset is relevant.

### Conflicting priorities

Instructions written by different teams or at different times may disagree about tooling, style, ownership, or release flow.

### Stale authority

Because the instruction file appears authoritative, an obsolete rule may override evidence available in the current repository.

### Review displacement

A detailed instruction file can make generated work look governed even when nobody has verified whether the instructions themselves remain correct.

More context can therefore reduce uncertainty for the agent while increasing risk for the organization.

## The Goal Is Minimum Sufficient Context

Minimum sufficient context is not the shortest possible instruction file.

It is the smallest set of current, authoritative constraints that the agent cannot reliably discover for itself.

Good repository context usually answers five questions.

### 1. What must never happen?

Include high-consequence constraints such as secret-handling rules, forbidden files, environment separation, destructive-operation limits, and human approval gates.

### 2. What is the real validation path?

Provide exact commands for tests, builds, linting, and behavior checks. Prefer executable commands over prose descriptions.

### 3. Where is the source of truth?

Point to the authoritative configuration or documentation instead of copying large sections into several files.

### 4. Which instructions apply here?

Use scoped instructions for specific directories, components, or workflows. A deployment procedure does not need to occupy the context of every documentation edit.

### 5. When should the agent stop?

Define conditions that require clarification or human action: unexpected working-tree changes, branch divergence, missing credentials, failed safety checks, changed PR scope, or production approval.

The best instruction often does not tell the agent everything.

It tells the agent what to inspect, what to verify, what not to assume, and where human responsibility begins.

## Context Should Be Layered, Not Dumped

A useful context system separates information by scope.

Root-level instructions should contain stable repository-wide rules.

Component-level instructions should contain only constraints relevant to that component.

Reusable skills or procedural references should describe specialized workflows such as publishing, migration, deployment, or rollback.

Executable tooling should enforce formatting, validation, and policy wherever possible.

The immediate task prompt should contain the current objective, acceptance criteria, and relevant exceptions.

This separation reduces the chance that every task carries the full weight of the organization’s operational history.

It also makes context easier to maintain.

When a deployment process changes, the deployment procedure can be updated without rewriting general coding instructions. When a lint rule changes, the linter remains the source of truth. When a temporary migration ends, the scoped migration guidance can be removed.

Context becomes modular infrastructure instead of an ever-growing prompt.

## Treat Agent Instructions Like Code

If repository instructions influence production changes, they deserve engineering discipline.

They need:

- a clear owner;
- review when modified;
- links to authoritative sources;
- explicit scope;
- removal of obsolete rules;
- validation against current repository behavior;
- and periodic deletion, not only accumulation.

Teams should also examine instruction changes as interventions.

If an `AGENTS.md` update adds 200 lines, what problem is expected to improve?

Which task category should benefit?

What would count as evidence that the change worked?

Could the rule be implemented as a test instead?

When will the instruction be reviewed or removed?

Without those questions, instruction files tend to grow through incident memory. Every failure adds a paragraph, while almost nothing creates pressure to subtract one.

## Lower Token Use Is Not the Same as Better Engineering

The research also exposes a measurement trap.

If instructions reduce token consumption, that may indicate better navigation.

It may also indicate that the agent stopped exploring too early.

If instructions increase token consumption, that may represent waste.

It may also represent necessary testing for a high-risk change.

Task success alone is not enough either. A patch can pass an immediate test and still produce additional review work, rework, duplication, maintenance cost, or short-lived code.

Context quality therefore cannot be evaluated through one number.

Teams need to connect the cost of agent execution with what happens after the first answer:

- Was the change accepted?
- How much human correction did it require?
- Did it create follow-up work?
- Did it survive later changes?
- Did performance differ by work category?
- Did the result improve enough to justify the added context and cost?

The value of context is downstream.

## Where GitMe Fits

GitMe does not decide what belongs in an `AGENTS.md` file.

It also does not prove that an instruction change caused a productivity improvement.

What it can provide is an outcome layer for examining what happened around changes in AI-assisted engineering work.

AI Effort Share shows where meaningful AI participation is concentrated.

AI Leverage keeps the question of leverage separate from simple AI adoption.

Work categorization helps distinguish feature development, bug fixing, refactoring, testing, documentation, configuration, security, and other forms of work.

Rework and historical comparison help reveal what followed the first implementation.

Effort Survival Cohort adds the durability dimension by showing how much past modeled engineering effort remains effective as the codebase evolves.

Together, these signals help teams ask a better question:

Did the new context system produce more durable engineering value with less corrective work, or did it merely increase the amount of agent activity?

That distinction matters.

A longer context file may make the agent appear busier and more compliant. More generated code may make the team appear faster. Neither demonstrates leverage on its own.

## Run a Context Experiment, Not a Context Rewrite

Teams do not need to redesign their entire instruction system at once.

A smaller experiment is more informative.

1. Select recurring task categories with enough historical examples.
2. Record the current instruction set and validation commands.
3. Remove duplicated, discoverable, obsolete, and task-irrelevant guidance.
4. Move specialized procedures into scoped skills or references.
5. Keep safety constraints and human approval boundaries explicit.
6. Compare similar work before and after the change.
7. Examine execution cost, acceptance, correction, rework, work mix, leverage, and durability together.
8. Restore removed guidance if the evidence shows that it prevented meaningful failures.

This is not a benchmark for ranking developers.

It is a way to evaluate whether the engineering system surrounding coding agents is improving.

## Less Context Requires Better Context

“Less context” should not become another slogan.

Removing a critical environment warning to save tokens is not optimization.

Hiding validation requirements from the agent is not simplicity.

Relying on the model to infer an unusual production constraint is not autonomy.

The objective is not less information at any cost.

The objective is less irrelevant information, less duplication, less staleness, and clearer authority.

Coding agents need enough context to act safely and effectively. Beyond that point, every instruction should justify the attention it consumes and the behavior it creates.

The best repository context is not the file that explains everything.

It is the system that gives the agent the right constraint at the moment it becomes relevant—and lets the team measure whether the result was actually better.

## Related GitMe Reading

- [AI Usage Is Not AI Leverage](https://www.gitdotme.com/blog-ai-usage-is-not-ai-leverage)
- [Why One AI Productivity Number Is Never Enough](https://www.gitdotme.com/blog-why-one-ai-productivity-number-is-never-enough)
- [When AI Reviews AI, Who Owns the Merge?](https://www.gitdotme.com/blog-when-ai-reviews-ai-who-owns-the-merge)
- [The Hidden Maintenance Tax of AI-Assisted Coding](https://www.gitdotme.com/blog-hidden-maintenance-tax-ai-assisted-coding)

## Sources

- [AGENTS.md — An Open Format for Guiding Coding Agents](https://agents.md/)
- [Gloaguen et al. — Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?](https://arxiv.org/abs/2602.11988)
- [Lulla et al. — On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents](https://arxiv.org/abs/2601.20404)
- [dos Santos et al. — Configuration Smells in AGENTS.md Files: Common Mistakes in Configuring Coding Agents](https://arxiv.org/abs/2606.15828)
- [GitHub Docs — Adding Repository Custom Instructions for GitHub Copilot](https://docs.github.com/en/copilot/how-tos/copilot-on-github/customize-copilot/add-custom-instructions/add-repository-instructions)

## Measure What Happens After AI Produces the Code

GitMe helps engineering leaders connect AI participation with modeled engineering effort, leverage, work categories, rework, historical comparison, and durability—without reducing engineering performance to activity counts.

[Start with GitMe](https://panel.gitdotme.com/signup)
