What is context engineering for AI coding agents?
Summary
- Context engineering decides what an AI coding agent can see before it acts: what enters the window, when it's retrieved, and what it can't reach.
- More context often makes agents worse, because scaffolding and bulk-loaded tools crowd out the signal the model needs to act on.
- Agents need four kinds of context: the code, the intent behind the change, the cross-system record, and the boundary of what they may not touch.
- Measure distribution, rework rate, review load, and time to merge, not license counts or token spend.
What is context engineering for AI coding agents?
A practitioner guide to context engineering for AI coding agents.
Context engineering is the practice of deciding what an AI coding agent can see before it acts. It governs what enters the context window, when it gets retrieved, how it resets between steps, and which systems the agent cannot reach. Prompting is what you say to a model. Context engineering is what the model can see when you say it.
The prompt gets most of the attention. The context, which does the actual work, usually gets whatever is left in the window. For an engineering leader, that's the difference between a few engineers getting strong results from coding agents and the whole organization getting them.
How does context engineering work?
Context engineering works by subtraction, not accumulation. An agent performs best inside a small, task-specific window, so the discipline is isolating context per phase, resetting it between steps, and retrieving information at the moment of use rather than loading it upfront. The constraint being managed is relevance, not capacity, and treating it as a capacity problem is the single most common mistake teams make when they set out to fix agent performance.
The clearest practitioner account of this we've heard is from Dex Horthy, founder of HumanLayer, on Dev Interrupted. He frames it as a smart zone and a dumb zone, and argues the entire job of building with agents is staying in the first one.
"The only thing that really matters is how do you optimize for staying in the smart zone."
— Dex Horthy, founder at HumanLayer, on Dev Interrupted, Dex Horthy on Ralph, RPI, and escaping the "Dumb Zone"
Horthy's team builds this in structurally rather than relying on prompt discipline, which is the right instinct: a parent agent owns the plan, each phase gets shelled out to a sub-agent so context stays isolated, a cheaper model writes while a stronger model checks, and context resets before the next phase begins. Nobody manages a window by hand, which is exactly the point.
The objection that smarter models make this obsolete doesn't hold up. The frontier moves without closing. Good AI products get built by finding the task sitting right at the edge of what a model can do, one it gets right some of the time, then engineering the context until it gets it right consistently. As models improve, that edge relocates to harder work. It doesn't disappear.
Why does more context make agents worse?
More context degrades agent output for a specific, mechanical reason: it crowds out relevance. A model weighs tokens in its window unevenly, and a window padded with scaffolding, bulk-loaded tools, or unfiltered records dilutes the signal the model actually needs to act on. The instinct when agent performance drops is to add context. The right move is usually the opposite.
Clare Liguori's team at AWS lived this the hard way. Liguori, senior principal engineer and a core maintainer of MCP, watched her team build heavy scaffolding around older models to make them behave reliably. Then a stronger model arrived, and the scaffolding that had been holding the old model together became the thing holding the new one back.
"We would have built a bunch of scaffolding, and then Sonnet 3.5 came out, and we were making the model actively worse because of all of the scaffolding on top."
— Clare Liguori, senior principal software engineer at AWS, on Dev Interrupted, Agent, skill, or MCP? Which to use and when to use them
The pattern isn't limited to hand-built scaffolding. It shows up in tooling design, too. Michel Tricot, CEO and co-founder of Airbyte, points to what happened when MCP first landed: clients loaded every available tool on connect and bloated the window before any work began. The fix wasn't a bigger window. It was a smarter client that could search for the tools a task actually required.
Liguori's own conclusion is the one we'd lead with: "the context is what makes all of this work." Not the harness around it. LinearB's write-up of that episode, Simplicity in model-driven agents and MCP drives real engineering gains, goes deeper on the scaffolding argument.
What kinds of context does an AI coding agent need?
We think most context-engineering efforts stop at one bucket when they need four: the code, the intent behind the change, the cross-system record that links the same person or feature across tools, and the boundary of what the agent may not touch. An IDE supplies the first for free, which is exactly why teams stop there. The other three have to be built deliberately, and skipping them is the most common gap we see.
1. The code
Relevant files, existing patterns, test structure. The category every team starts with, and the only one that doesn't require any new engineering to get right.
2. The intent
What the change is for, what is explicitly out of scope, which architectural option was chosen and why. Horthy's team builds this as a design discussion document generated before implementation: a hundred to two hundred lines of markdown covering the desired end state, where things stand now, what is out of scope, the relevant patterns in the codebase, and a short list of open design questions. Small enough to read in full. Cheap enough to throw away.
3. The cross-system record
An agent working a real delivery problem has to know that a person or a feature in one system is the same person or feature in another, and most agentic tooling is not built to answer that question. Tricot frames the underlying task as entity resolution.
"Every data problem that an agent has is just a search problem."
— Michel Tricot, CEO and co-founder at Airbyte, on Dev Interrupted, Your agents are starving! Airbyte's Michel Tricot on the data ingestion crisis
Without that linkage, an agent cannot follow a feature from a ticket through review into production, and it will confidently produce work that ignores context it never had a way to find.
4. The boundary
What the agent may not reach, and whether it knows that. We think the second half of that is where most implementations fail: walling an agent off from a system with no way to signal what's missing just pushes the agent to improvise around the gap. Tricot's version of the fix is keeping metadata about blocked systems visible even when the underlying data is not, so an agent can recognize the gap and request access rather than guess.
What are examples of context engineering for AI coding agents?
Four patterns carry most of the value, and they all follow from the same principle: give the agent only what the current step requires, and rebuild that set constantly rather than accumulating it.
Reset early and often. Isolate context per phase instead of carrying one long conversation forward. Sub-agents are the mechanism for that isolation. The goal is a clean window at the start of every phase; multi-agent architecture is the means, and a clean window is the point.
Front-load alignment into a cheap artifact. Produce the intermediate document before any code gets written. This is the last point where correcting course costs minutes instead of a rewrite, and it gives a human reviewer something legible to check rather than a diff.
Disclose progressively. Liguori's team at AWS serves roughly 16,000 APIs through one remote MCP server. The server surfaces curated skills that describe how to use those APIs for a specific job, rather than exposing all 16,000 tools directly. The agent pulls an instruction only when the task calls for it. That's the model worth copying: capability behind a search, not capability dumped into the window.
Retrieve just in time. Discover step by step rather than in bulk. Asked for the lifecycle of a customer on a feature, an agent should identify the feature, then the ticket it maps to, then the conversation that mentions it, keeping the window lean the whole way through.
What are the most common context engineering mistakes?
The most expensive mistake is a category error: treating context engineering as a technique when it has to function as a system. Context engineering reliably works for the engineer already immersed in AI tooling and reliably breaks down when handed to a hundred people who aren't, and most rollouts never diagnose why.
We think the reason is that the practice usually lives in one person's undocumented intuition rather than in the product itself. Horthy's account of building RPI, research-plan-implement, is the clearest evidence for this we've seen. It worked. Engineers already deep in AI tooling picked it up and got real results. Then those engineers tried to hand it to their teams, and it stopped working.
"Most people who hadn't been obsessed with AI content and following all the stuff and in all the group chats just couldn't get good results."
— Dex Horthy, founder at HumanLayer, on Dev Interrupted, Dex Horthy on Ralph, RPI, and escaping the "Dumb Zone"
Horthy traces the failure to something specific: he'd catch himself telling workshops to sprinkle certain magic words into the planning step or results would degrade. That's the tell. A practice that depends on incantations is one person's habit that happens to work, not a practice a team can inherit, and the fix belongs in the product rather than in the prompt. His framing since has been to do the context engineering for people, so the default path produces a usable result without anyone needing deep intuition about a specific model.
This is exactly why we think AI tool adoption numbers are the wrong thing to track. A license count tells you the tool is installed. It says nothing about whether the practice that makes the tool work has spread past the handful of engineers who were going to figure it out on their own. The distance between those engineers and everyone else holding the same license is the real state of a rollout, and it's a gap adoption dashboards are structurally unable to see. LinearB's own analysis of that split, Your software factory needs a context layer, measures the same pattern across hundreds of organizations.
There's a third mistake sitting underneath the first two, and it's a governance one, not a technical one: treating AI-generated code as exempt from the review standard applied to everything else. Liguori's teams hold a firm line here.
"You have to take accountability for the code that you produced, even if you generated it using a model or if you wrote it by hand."
— Clare Liguori, senior principal software engineer at AWS, on Dev Interrupted, Agent, skill, or MCP? Which to use and when to use them
Read it before it becomes a pull request. Context engineering without that rule doesn't make a team faster. It just produces slop faster.
How do you measure whether context engineering is working?
Measure outputs, not inputs. Token spend and license counts prove that money moved and nothing else; they say nothing about whether the practice is compounding or whether it's about to plateau. This is the question we think most AI measurement dashboards get backwards, and it's the reason a maxed-out usage chart can sit next to a stalled delivery pipeline without anyone noticing the contradiction.
Horthy's version of the test is blunt, and it matches what we've seen in the data:
"It's not about the inputs, it's about the outputs."
— Dex Horthy, founder at HumanLayer, on Dev Interrupted, Dex Horthy on Ralph, RPI, and escaping the "Dumb Zone"
Tricot makes the same case from the finance side, and rejects the standard six-month ROI window outright. He compares it to the move to cloud, where a lift-and-shift produced a worse bill until teams actually rebuilt for the new model. Judge context engineering on whether the practice is still improving, not on a quarterly snapshot.
In practice, that means watching four things:
- Distribution rather than adoption. How many engineers are getting agent-assisted changes merged, against how many hold a license.
- Rework rate on agent-assisted changes, compared with changes made without them.
- Review load. If context engineering is working, the volume reaching human reviewers gets smaller and more substantive, not larger.
- Time from intent to merged change, which is where better context should show up before anything else.
This is precisely the correlation LinearB was built to surface: AI activity mapped against pull requests and delivery outcomes, so a rollout's distribution and rework rate stop being guesses. We benchmark that against 8.1 million pull requests from 4,800 teams in 42 countries, which is the scale you need before that claim is more than a hope.
Frequently asked questions
Is context engineering different from prompt engineering?
Yes. Prompt engineering concerns how a request is phrased. Context engineering concerns what the model can see when it receives that request: retrieval, isolation, resets, and boundaries. One is writing. The other is systems design.
Does context engineering stop mattering as models improve?
No. Better models raise the ceiling on what an agent can attempt, which moves the hard tasks to a new edge. The practitioners getting outsized results from each model generation are the ones still doing this work.
What should a team change first?
Generate a short design document before implementation and read it. It's cheap, it catches misalignment while correcting is nearly free, and it gives a reviewer something legible.
Why do a few engineers get good results and the rest get nothing?
Because the practice is living in individual intuition rather than in tooling and defaults. The fix is to encode the practice, not to train harder.
Does a bigger context window remove the need for this?
No. A larger window raises the ceiling on what fits, and relevance still decides what the model does with it. Bulk-loading a large window just reproduces the same failure at greater cost.