Scott Wueschinski
← All articles

Insight

Everyone Is Reading the Wrong Document About Agents

The spec, the demo, and the model are the least interesting things in production. The boring layer underneath decides who ships and who stalls.

· 9 min read

There is a habit I keep running into from the forward deployed seat, and it costs teams more than any model choice they will ever make. When a new capability arrives, everyone crowds around the most legible artifact in the room and studies it as if that is where the outcome gets decided. They read the specification. They watch the demo. They benchmark the model. Then production happens, and the thing that breaks is never the artifact they were staring at. It is the layer underneath, the one nobody put on a slide.

This is the single pattern connecting the standards fight around agent tooling, the eval work most teams skip, and the demos that collapse the moment real traffic hits. Three different problems on the surface. One mistake underneath. People keep grading the visible thing and ignoring the thing that actually determines whether the system survives a Tuesday.

The visible artifact is the decoy

Consider how the Model Context Protocol conversation plays out inside most engineering orgs. The spec gets circulated, someone summarizes the primitives, a working group forms to decide whether the wire format is mature enough. All of that energy points at the one document that does not matter. As I argued in MCP is a distribution play, not a protocol, the specification is JSON-RPC with some primitives bolted on: useful, clean, and boring. The SDK was downloaded roughly 100,000 times in its first month, and by March 2026 it had reached 97 million monthly downloads, a 970x jump in eighteen months. The React npm package took about three years to reach 100 million monthly downloads. MCP got to comparable scale in sixteen. No schema wins that fast. Defaults do. The real question was never how the transport works. It was who owns the tool surface an agent reaches for when a user says “check my calendar” or “pull the invoice.”

Now hold that next to the demo problem. A team wraps a retry loop around a single prompt, points it at a curated input, and calls the result autonomous. The stage version looks alive because the demo path is the only path anyone ever runs. But as I wrote in most agent demos are just a for-loop with anxiety, the property that separates a demo from a system is architectural, and it hides three questions the audience never gets to ask: does the thing hold state, does the model actually choose its own tools, and can the whole system be evaluated. Leading agents complete only 30 to 35 percent of multi-step tasks reliably, and that is a compounding math problem the demo is designed to conceal. In a six-step workflow where each step succeeds 95 percent of the time, the compounded success rate drops to 74 percent. The demo shows you step three going well. It never shows you the sequence.

And the eval problem is the same misdirection wearing a lab coat. The agent worked in the notebook, passed the demo, and got leadership sign-off. Then, as the incident in the eval you skip is the outage you schedule describes, a tool call started returning malformed JSON within two days and the agent silently continued on bad data. The failure was a case nobody wrote a test for. Nearly every agent incident I have worked traces to a failure mode somebody looked at, judged too rare to matter, and skipped.

Three surfaces. Three decoys. The spec, the happy-path demo, and the green checkmark all share one property: they are the parts of the system that photograph well and reveal nothing about how it behaves under load.

Why smart teams keep grading the decoy

This is an incentive structure, and it is worth naming precisely so you can fight it..

The visible artifact is legible. A specification can be read in an afternoon and discussed in a meeting. A demo can be scheduled, rehearsed, and approved. A model benchmark produces a number you can put on a chart. The boring layer underneath has none of those properties. Distribution, state management, and evaluation are invisible until real traffic exposes them, so they never generate the same organizational pull. Nobody gets a round of applause for the registry policy that decides which tool an agent trusts, or for the checkpoint logic that lets a workflow resume mid-sequence. Those things only get noticed when they are missing.

Scope is the second driver, and it is the one I hear most often when I ask why a team shipped without evals. A real eval harness feels enormous, so teams treat it as all or nothing: either we build the full pipeline or we ship on judgment. Almost everyone ships on judgment. The industry numbers confirm the split. LangChain’s 2026 survey of more than 1,300 practitioners found that nearly 89 percent have implemented observability for their agents while evals adoption sits at 52 percent. Observability is the easier artifact to adopt because it tells you what already happened. Evals are harder because they force you to imagine what has not happened yet.

The third driver is the most dangerous, because it feels like progress. Agent failures return valid, well-formed responses that are semantically wrong.. A tool call succeeds against an outdated schema. A retrieval pulls from the wrong context window. A planning step executes flawlessly and solves the wrong subproblem. Nothing crashes. Nothing alerts. From the outside the system looks healthy. So a team that grades the decoy gets rewarded with a clean dashboard right up until a customer discovers the damage. The reliability research is blunt about the mechanism: frontier models stay capable on individual steps but fall apart across the sequence, succeeding on tasks that take a human expert a few minutes and degrading sharply as tasks stretch to hours. The model was fine. The system could not hold it together, and no single-turn accuracy metric will ever surface that.

Put those three together and you get an organization that is optimizing exactly the wrong thing with total sincerity. It reads the spec because the spec is readable. It demos the happy path because the happy path demos well. It buys observability because observability installs cleanly. Every one of those choices is locally rational and collectively fatal.

The bill arrives quietly, and it compounds

The reason this pattern is so expensive is that the cost of doing nothing accrues in the dark, and by the time it is visible it has already compounded..

On the distribution side, the cost is losing the default. Once an agent is trained, prompted, and wired to reach for someone else’s server in your category, you are fighting muscle memory at the model layer.. A competitor who ships a technically superior tool two quarters later does not get a second look, because the agent already reached for the incumbent. OpenAI, Google, Microsoft, and Salesforce all shipped MCP support within thirteen months, and the vendors racing to build servers are doing it for placement.. Defaults compound the same way shelf space in a store compounds. The team still auditing whether the protocol is mature enough is paying, every quarter, for a position it does not hold.

On the reliability side, the cost is the silent wrong answer that a customer finds before your dashboard does. A user gets bad billing guidance. A memo cites a fabricated policy. A retry loop burns tokens re-running a broken step against broken context and calls it resilience. LangChain’s respondents named quality as the production killer, with 32 percent citing it as a top barrier, and that number is the visible tip of a much larger accrual. The public incidents make the shape of the risk concrete: an AI assistant deleted a production database despite instructions forbidding it, an autonomous operator made an unauthorized purchase that bypassed user confirmation, and a government chatbot handed out illegal business advice. None of those were model failures. Each was a boundary that lived in a prompt instead of in a test.

This is why pilots pile up while scale refuses to follow, and why boards start canceling programs that demoed beautifully. The organization priced the model and the demo, and it never priced the compounding. A workflow that fails 26 percent of the time produces well-formed wrong outputs that erode trust one interaction at a time until someone senior asks why the numbers do not reconcile.. The cost of doing nothing on the boring layer is a debt that accrues interest in silence..

Build the operating model around evidence

The fix is an operating discipline that treats the invisible layer as the actual product and forces it to produce evidence.. Three moves make it real.

First, audit the surface instead of the spec. For anything touching agent tooling, ask three concrete questions. When your agent needs a tool in your category, whose server does it reach for first. Who controls the registry your enterprise buyers trust. And what does it cost you, per quarter, to not be the default. Anthropic’s 2026 roadmap tells you where the leverage is, because it points at enterprise authentication, multi-agent coordination, and a curated, verified server registry with security ratings. A rated directory is a distribution channel with an editorial policy. Whoever runs it decides which servers surface first and which get buried. Enterprise buyers already sense this, which is why they demand registry, approval, and audit layers before granting broad write access. Own that surface inside your own four walls before someone else defines it for you.

Second, ship the three evals that cover most of your real risk this week, not the six-month harness you will never finish. The tool contract eval feeds the agent malformed JSON, a partial payload, a timeout, and an empty result, then confirms it halts or routes to a fallback rather than reasoning on garbage. Bad input is Tuesday.. The boundary eval hands the agent a task that tempts it across its permission line and proves it stops at the gate, because a test that demonstrates refusal is a boundary and an instruction in a prompt is only a wish. The multi-turn drift eval runs a realistic ten-turn conversation and checks whether turn ten still serves the objective from turn one, since a wrong plan executed perfectly is still a wrong plan. That is roughly a day of engineering, and it maps directly onto the three failure classes that produce production headlines.

Third, make the system prove its own state and tool discipline. If it cannot checkpoint progress and resume mid-sequence, it is a goldfish with API access, and no amount of retry logic will change that. Let the model choose tools inside guardrails that refuse everything outside its lane. Instrument every decision so you can reconstruct what happened and why. Then close the loop: let failing production traces flow back into the offline eval set so the whole system gets harder to fool over time. The eval you skip is a scheduled outage where you have chosen only the timing and whether you or a customer discovers it first..

The Tuesday test

The through line across distribution, reliability, and architecture is a single standard you can apply to any agent program right now. Stop asking whether the demo impressed the room, whether the spec looks clean, and whether the benchmark ticked up. Ask whether the system can prove, on an ordinary Tuesday, under real load, exactly why it did what it did and exactly what it refused to do.

That standard reorganizes everything. It turns the registry question from a standards debate into a placement strategy. It turns the eval backlog from an all-or-nothing project into three tests you write before Friday. It turns the demo from a performance into a claim you can defend with traces. The teams that win in 2027 will be the ones who built the boring layer into the product from day one, who own the defaults in their category, and who can stand up under load and show their work.. The visible artifact was always the decoy. Build for the layer underneath it now, because that is where the next two years actually get decided.