Scott Wueschinski
← All articles

Insight

Production agents die on the constraints teams skip before the demo

The demo always lands. What breaks in production is the boring constraint work: budgets, context, tool surfaces and architecture. Here is the operating model that fixes it.

· 9 min read

From the forward deployed seat, I show up after the demo has been sold and before the pager starts going off. That gap is where the same story keeps playing out, in four different costumes. A team ships an agent. It looks great in the room. Three weeks later it is stitching nine tool calls to do the work of two, quietly duplicating records, burning an annual budget by April, or confidently operating on a version of the business that stopped being true a month ago. Every time, the team’s first instinct is to reach for a bigger model, a smarter prompt, more tools, or more agents.

That instinct is the disease, not the cure. The failures I see in production are almost never failures of reasoning. They are failures of constraint. Agents fail because nobody bounded what they can spend, curated what they can call, engineered what they can see, or tested whether the architecture the org chart suggested actually pays off under load. The reasoning layer gets all the attention because it demos well. The constraint layer gets skipped because it does not, and the constraint layer is the entire difference between a system that compounds and one that decays.

The demo is a liar, and the divergence is silent

Here is the uncomfortable truth that ties every one of these failure modes together: a good agent and a bad one look identical in the demo. A decaying MCP server and a compounding one both land the same happy-path task in front of the same impressed room. A single agent and a four-agent swarm both produce a plausible answer on the slide. An agent with a hard budget and one with no ceiling both cost roughly the same for the first ten minutes. The retrieval that stuffs the top ten near-duplicate chunks into the window looks fine on the question you rehearsed.

The divergence happens in production, and it is a function of work you either did or skipped before the demo. It is slow, and it is quiet, which is exactly why it survives review cycles. Skip the design of intent-level tools and you pay a tax every time the agent stitches six primitives together, bloats the context, and lets a retry duplicate a record that nobody notices until someone checks the database on a Thursday. Skip the budget primitive and you pay when four agents fall into an infinite loop, run for 11 days straight in November 2025, and empty a wallet while an 80 percent alert fires into a void. Skip context engineering and you pay in rework and trust withdrawals every time the agent invents an answer because a tool call timed out.

None of these show up as a line item this quarter. That is what makes them dangerous. The Cost of Doing Nothing is the compounding tax you pay on every request because you optimized the visible layer and ignored the one that governs behavior under real load..

Why smart teams reach for the wrong lever

The reason this pattern repeats is that the wrong lever is always the more legible one. Upgrading a model is a purchase order. Adding agents is a diagram you can present. Wrapping your REST API one-to-one and exposing forty tools feels like thoroughness. Wiring a Slack alert feels like governance. Each of these is easy to justify, easy to ship, and easy to mistake for progress.

I have watched the model-upgrade version of this loop run for two years, and the script never changes. An agent breaks, the diagnosis is instant and unanimous that the model is not smart enough, the team moves to the newest frontier model, sees a marginal bump, and hits the same wall in a slightly different shape three weeks later. The New Stack named the real constraint this summer: the bottleneck for AI agents is the context layer. What actually breaks in the field clusters into three boring categories, and none of them are fixed by better weights. Stale retrieval, where the vector store is behind reality. Missing tool results, where a fetch fails and the agent guesses instead of stopping. Context bloat, where the window fills with boilerplate and near-duplicate chunks before the real task even loads. Chroma Research made the last point brutally concrete when they tested eighteen frontier models and found every one degraded as input length grew, with a 200K window showing serious loss by 50K tokens. A bigger window is not an escape hatch when the retrieval feeding it is read every file and hope. I make the full case for this in the model is not your bottleneck, your context is.

The multi-agent version of the same mistake trades the model lever for the architecture lever. A team has a single agent that works, then splits it into a planner, a researcher, a writer, and a critic because that felt more serious. What they bought was an org-chart metaphor, and they pay for it in latency, failure surface, and debugging nights on every request. Google Research’s agent scaling study is the finding worth taping to your monitor: multi-agent coordination delivers +81% improvement on parallelizable tasks but causes up to 70% degradation on sequential ones. Same technique, opposite outcomes, and the deciding variable is the shape of the work, not the cleverness of the design. Most real business workflows are stateful and ordered, which means fanning them out to a swarm introduces handoffs where none existed and lets errors cascade instead of cancel. A four-agent essay grader makes the cost visible: four LLM calls per essay, roughly 4x the cost and latency, for a 4 percent relative accuracy bump. I unpack when the coordination tax is worth it in multi-agent is a coordination tax you may not need.

The through-line is that every one of these legible levers lets a team feel like it is solving the problem while leaving the actual constraint untouched. The model upgrade dodges the context work. The swarm dodges the discipline of making one agent excellent. The forty-tool API mirror dodges the design of intent-level capabilities. The dashboard dodges enforcement. Legibility is the trap.

What the skipped work actually costs

The economics of skipping the constraint layer are worse than they look, because the token math is not linear. Everyone watches token prices fall and assumes the cost problem solves itself. Agents make many calls. They loop, plan, call tools, retry, hand off, and resend their full context at every step, so by step twenty you have paid for the same history twenty times.. That architecture is a multiplier: agentic workflows consume 5 to 30 times more tokens per task than a standard chatbot query. TechCrunch put the trend line on it with per-developer consumption rising about 18.6x in nine months.

This is how even disciplined organizations get blindsided. At Uber, Claude Code adoption jumped from 32% to 84% of a 5,000-engineer org between December 2025 and March 2026, and by April the entire annual AI budget was gone. Their CTO said the quiet part out loud, that he was back to the drawing board because the budget he thought he would need was already blown away. That is what an agent shipped without a budget primitive does at scale, and it is why forecasting fails by default: 85% of companies miss AI cost forecasts by more than 10%, and nearly 25% underestimate by 50% or more, largely because instrumentation stops at the model API level rather than at the individual agent.

Now stack the other taxes on top. The context tax shows up as inference bills growing faster than the value the agents provide, because upgrading the model while the context is broken fails more expensively, paying frontier prices to feed garbage into a bigger window. The architecture tax shows up as a swarm that hallucinates in four voices instead of one, which is a bigger incident with a nicer diagram. The tool-sprawl tax shows up as bloated context and an agent that cannot reliably pick the right call, a bill that compounds silently for months.

The market has already priced this in. Gartner predicts over 40% of agentic AI projects will be canceled by 2027 due to escalating costs and inadequate controls. The teams that survive that cull will be the ones whose agents cannot spend a dollar they were not given, cannot call a tool they were not curated for, and cannot see context that was not engineered to be fresh.

The operating model: constraints as first-class primitives

The fix across all four failure modes is a single discipline stated four ways. Treat the constraint as a primitive the agent cannot run without, set before production rather than after the postmortem, and prove it under load against a baseline you actually built. Here is what that looks like in practice.

Bound the spend at the infrastructure layer. Give the agent a wallet before you give it a personality. That means a per-task token ceiling and a hard step limit that are non-negotiable, a cost circuit breaker in the request path that evaluates spend rate before each call, and a safe-mode fallback that narrows execution to read-only tools rather than crashing. The cap has to live in the infrastructure, because application-layer cost counters reset on process restart and can be bypassed by exception handling. If your budget lives inside a try/except block, you have a suggestion. The full argument is in your agent needs a budget, not just a prompt, and the sequence is deliberate: budget it cannot exceed first, prompt second.

Curate the tool surface around outcomes. Do not open your API docs first. Write the three to five tasks an agent must complete against your system, then design the smallest surface that covers that work, one tool per outcome where you can manage it. A single create_issue_from_thread tool beats get_thread plus parse_messages plus create_issue plus link_attachment, because intent-level tools mean the agent does the work in two calls instead of nine. Keep the count honest at roughly 5 to 8 tools per server; past that you are back to wrapping endpoints. Make every call idempotent with client-generated request IDs so a retry cannot corrupt the database. This is a three-day sequence, not a three-sprint project, and if you cannot do it in three days your scope is wrong. The playbook is in the three-day MCP server design playbook.

Engineer what the model sees as a system. Version your retrieval so freshness is a guarantee rather than a hope. Instrument tool calls so a failed fetch stops the agent instead of triggering a confident guess. Budget the window ruthlessly and stay out of the dumb zone that Dex Horthy identified, where performance degrades once more than roughly forty percent of the context window is consumed. Rank and dedupe what you inject, because ten chunks from the same document section waste nine slots on near-duplicates and hand the model one perspective repeated ten times.

Earn the architecture with numbers. Begin with a single-agent prototype to establish baseline capabilities, and transition to multi-agent only when testing reveals limitations that single-agent optimization cannot resolve. The word that matters is measurable. Distinct roles like planner, reviewer, and executor might suggest multiple agents, but a single agent can wear four hats without four bills. Reserve the swarm for work that is genuinely parallel, multi-role, or larger than one context window.

Build the constraint layer this quarter

The common failure of the next twelve months will look exactly like the common failure of the last two years, dressed in newer models and bigger context windows. Teams will keep optimizing the visible layer because it demos, and keep paying the silent tax because the constraint layer does not.

The teams that win the next cycle will invert that. They will treat spend, context, tool surfaces, and architecture as design constraints proven under load, not features bolted on after the invoice lands. Every one of those constraints is something you own today. It sits in the retrieval you version, the tool surface you curate, the budget you enforce at the infrastructure layer, and the single agent you make excellent before you ever earn a second. The next capability jump for your agents is in the boring constraint work everyone skips because it never shows up in the demo.. Go build it, starting Monday.