Scott Wueschinski
← All AI and Agentic POV

Pilots do not scale because nobody owns the exception queue

Agents nail the happy path. The exceptions land on a team nobody staffed, and that is where the pilot quietly dies. Exception ownership is the scaling constraint.

Agentic Retail POV agentic-operations

· 3 min read · Source: AlignX AI, Designing Human-in-the-Loop for Agentic Workflows ↗

Every agent pilot I have reviewed in the last eighteen months died the same way. Not in the model. Not in the demo. In the exception queue that nobody agreed to own.

The pattern is boringly consistent. An agent gets built for a well-scoped retail workflow: returns triage, vendor dispute resolution, promo exception approvals, supplier onboarding. It handles the happy path beautifully. In the sandbox it clears 80 percent of volume, the deck says “80 percent automation,” and the room nods. Then it meets production, and the 20 percent it cannot resolve lands somewhere. That somewhere is a team that was never staffed, never trained, and never told they now own the hardest fraction of the work.

That is where the pilot quietly dies. Not with a failure. With a backlog.

The constraint moved, and most business cases did not

For two years the industry has told itself that AI does not scale because the models were not good enough. That story is now wrong. The scaling gap is not primarily a technology problem. The models are capable. The tooling has improved dramatically. When 78 percent of enterprises are running a pilot and only a small fraction reach organization-wide use, the binding constraint is not intelligence. It is operations.

The clearest articulation of the trap I have read comes from a piece on designing human-in-the-loop workflows. The most common mistake in human-in-the-loop design is turning humans into quality assurance checkers. On the surface, this seems reasonable. In practice, it collapses under production load. The prescription is correct and everyone repeats it: agents should handle routine cases autonomously. Humans should handle exceptions.

Here is the part the board never hears. “Humans should handle exceptions” is not a design principle. It is a headcount line. And it belongs in the business case on day one, not discovered in month four when the escalations pile up.

Exceptions are not noise. They are the job you kept.

Think about what an exception actually is. It is a customer request that is technically outside policy but contextually reasonable. It is conflicting shipment data or a customs issue. It is a vendor change or a customer-facing compensation decision. These are exactly the cases the source calls out: a customer request is technically outside policy but contextually reasonable. The agent flags it. A human decides whether to make an exception.

Those are the highest-judgment, highest-stakes, highest-emotion interactions in the whole workflow. The agent skims off the easy volume and hands your team a concentrated stream of the hard stuff. If you staffed that team assuming the old case mix, you have quietly made their job worse. Average handle time goes up. Morale goes down. The people who understand the exceptions best are the first to leave. And the pilot’s promised savings never materialize because you traded high-volume easy work for low-volume expensive work without changing who catches it.

This is the Cost of Doing Nothing hiding in plain sight. The CODN framework usually points at the risk of not deploying. Flip it. There is also a Cost of Deploying Nothing Downstream: an agent that automates the front of a process while leaving the exception tail unowned does not reduce cost, it relocates and concentrates it, and then it stalls.

Put the headcount question in the business case, or do not fund the pilot

I now ask three questions before any agent gets funded past pilot. None of them are about the model.

First: what is the exception rate at full production volume, not demo volume? Demo exception rates are fiction. Real volume surfaces the long tail.

Second: who owns the queue, by name and by function? Not “the ops team.” A named owner with authority, context, and a defensible rationale. The governance literature is blunt that oversight requires trained people with the authority to intervene, and that presence is not the same as practice.

Third: what is the fully loaded cost of that ownership, and does the business case still clear its hurdle rate with that number in it? If the answer is no, you have not found a bad model. You have found a workflow that is not ready to automate.

The agents that survive a business case are the ones where someone signed up to own the tail before the pilot launched. The ones that become next year’s write-off are the ones where everyone admired the 80 percent and nobody budgeted for the 20.

Model quality is table stakes now. Exception ownership is the moat. Fund the queue, or do not fund the agent.