Home
/
Blog
/
Simplicity in model-driven agents and MCP drives real engineering gains

Simplicity in model-driven agents and MCP drives real engineering gains

Photo of Andrew Zigler
|
Blog_Simplicity_in_model_driven_agents_2400x1256_e67f24d00b

The hardest problem in agent development right now is not building. It is knowing what to leave out. Clare Liguori, Senior Principal Software Engineer at AWS, has spent the last three and a half years watching the field lurch from inline code completion to vibe coding to full agentic harnesses, and the lesson she keeps returning to is that the model, not the machinery around it, should carry the load.

As the technical lead on the open source Strands Agent SDK and a core maintainer of MCP, she sits at the center of two of the most consequential shifts in how engineering organizations ship AI. Both point in the same direction. The scaffolding that once felt like progress has become the thing holding teams back, and the fastest-moving teams are the ones tearing it down.

Model-driven agent development lets the model do the heavy lifting

Strands began inside the Agentic AI org because existing frameworks carried too much cognitive overhead. Getting something into production took six months. The team wanted a declarative interface that focused on the things that actually matter, the model, the system prompt, the tools, and stripped away everything else.

The timing mattered. The scaffolding built to make older models behave reliably turned into a liability the moment better models arrived. "We would have built a bunch of scaffolding, and then Sonnet 3.5 came out, and we were making the model actively worse because of all of the scaffolding on top," Liguori says, because that scaffolding stopped giving the model the right context. Sonnet 3.7 repeated the pattern. The capabilities jumped, and the harness fought the very reasoning it was meant to enable.

The instinct to build more is deeply human. Engineers form a personal connection to the systems they create, which makes those systems agonizing to tear down even when their entire lifetime has been six months. The model-driven approach asks teams to resist that pull and let the model do most of the work. Assume that in six months a single line change in the agent configuration swaps in a new model, and the new reasoning and tool-calling capabilities become available automatically. That assumption only holds when nobody has boxed the model in with a layer of constraining logic on top.

MCP protocol simplification finally makes remote servers realistic

The same discipline now shapes the MCP spec. The most significant change is how much easier it becomes to build remote servers over the HTTP transport. Standard IO servers have long been the barrier, and Liguori has been pushing to get out of that game entirely. "I think that really any SaaS provider should have an MCP endpoint for their APIs," she says. The trouble is that implementing anything beyond basic tools as a remote provider has been hard.

Elicitation is the clearest casualty. Responding to a user with a form or a URL to gather more information required the server to be heavily stateful, which meant building a streaming API into a web server, and the feedback on that difficulty came through loud and clear. The transport shift changes the math. The spec is moving to a stateless design with typical request-response patterns, which means capabilities like elicitation stop being impossible and start being routine.

Stability drives the rest of the design. The core spec stays deliberately quiet while experimental ideas live in extensions. Tasks, which model long-running jobs rather than short request-response tools, generated churn as the design evolved, so it moved out to an extension where breaking changes do no damage. MCP Apps, already supported in ChatGPT, lets a server provide a full UI widget through the same mechanism. Things graduate into the main spec only once real-world usage proves them out. The AWS MCP server for APIs shows what this simplification buys. One remote HTTP server exposes roughly 16,000 AWS APIs through curated skills owned by individual service teams, instead of hundreds of sprawling individual tools nobody can reason about.

Agentic coding assistants are replacing custom-built agents

Simplification also changes who should build agents at all. When coding assistants became powerful, the case for bespoke agents collapsed. Liguori also works on Kiro, AWS's agentic coding assistant, and internally the effect was dramatic. Thousands of engineers stopped building on-call troubleshooter agents from scratch and instead wrote a small custom agent config against Kiro. The QCLI, now the Kiro CLI, launched to public production in three weeks, a timeline that was previously unthinkable.

The productivity gains are real and measurable. Pilots with different teams are showing a four to five X increase, and the gains track new habits as much as new tools. Teams that see the jump have changed how they work, not just what they run.

The deeper pattern is organizational. Conway's Law has been applied to agents, so the number of agents a company ships maps directly to its org chart. Every team feels obligated to own one, and customers now describe fleets of 500 agents sprawling across internal teams. That is too many. Most products need one agent plus a set of tools and skills disclosed progressively, not dozens of one-off micro agents each carrying a single skill. Standardizing on a shared assistant, configured many ways, replaces the duplicated bespoke work that org charts keep generating.

AI code review automation catches the slop humans miss

Accountability is the practice that separates the teams pulling four to five X from the ones drowning in output. The old rite of passage was a new hire fumbling a Git command. The new one is an AI slop pull request, generated by a coding assistant and pushed by a human who never read it. The rule holds regardless of authorship. "You have to take accountability for the code that you produced, even if you generated it using a model or if you wrote it by hand, you have to take the same accountability," Liguori says. The assumption on her team is simple. You read it before it becomes a pull request.

Automated review handles the rest. Her teams run AI code review to catch the total slop, but also to catch the nitty-gritty questions humans should not have to spend attention on anymore. Is this maintainable? Did you add tests, and do they make sense? Those checks used to consume human reviewers who missed the forest for the trees, nitpicking style while never asking whether the change was the right thing to ship or whether it touched code too sensitive to break in production. Now the machine gets to be picky. "The AI code reviewer and your AI code generator get to chat with each other," she says, so most of that back-and-forth resolves before a human ever opens the diff.

At an organizational scale, the danger is fragmentation. If every team runs its own review prompts, enforcement splinters. AWS addresses this with internal mechanisms that install custom agents, MCP servers, and skills the way a package manager installs dependencies, plus hooks into the code review system itself so a custom agent can review a pull request and add comments directly. Standardizing those review prompts across the org keeps the quality bar consistent as the volume of code moving through the codebase climbs past anything a human could read.

The fastest teams build context, not constraints

The model-driven bet has now been tested well beyond the team that made it. Inside AWS, three major projects reached production in under six weeks each before Strands was ever open sourced, which is when the success stopped looking like luck and started looking like a repeatable method. The preview drew the same story from outside. One customer took a goal to route 25% of traffic through agents, picked up Strands, pushed an agent to production in a month, and hit the goal by mid-year. What generalizes is not the framework so much as the discipline behind it. Keep the surface small and let the model carry the load.

That discipline is also the way out of the organizational tangle. For 25 years engineering absorbed the gospel of service-oriented architecture, and now teams are shoving markdown files around, sometimes uploading them to Slack for a colleague. Some order has to emerge, and Liguori's answer is fewer, sharper building blocks. A tool that fetches context or takes an action. A skill that teaches a model how to use those tools. An MCP server, CLI, or API underneath. The question for most teams is not which agent to build but whether they need to own an agent at all, or just a skill.

Her closing advice to engineering leaders is the same lesson that opened the conversation. Always think about simplicity. The models are amazing now, and the work is building context around them rather than constraints on top of them, because "the context is what makes all of this work."

To hear more of Clare Liguori's insights on model-driven agent development, MCP protocol simplification, and AI code review automation, listen to the full episode on the Dev Interrupted podcast.

Headshot3_d7231cbda7

Andrew Zigler

Andrew Zigler is a GTM Engineer at LinearB and the host of Dev Interrupted, a twice-weekly podcast and newsletter where 40k+ builders decode the transition to AI-native development and agentic orchestration. A classicist by training with a degree from The University of Texas at Austin, Andrew spent his early career teaching in Japan before channeling his interdisciplinary instincts into the tech world. His polymath background informs everything he builds, from automated workflows to the stories he tells about the seismic shifts reshaping software creation.

Connect with

Your next read