Scott Wueschinski
← All articles

Insight

Agentic Retail Rewards the Workflow. Boards Keep Funding the Artifact.

Six failures in agentic retail share one root cause: the enterprise measures and funds the wrong unit of value. Fix the operating model, and the margin follows.

· 10 min read

Sit through enough retail and CPG board meetings this year and you start to hear the same sound underneath every agenda item. It is the sound of a well-run enterprise funding the wrong thing with total confidence. The catalog team invested in a PIM and believes the data is clean. Procurement grinds a discount on a three-year seat term. The governance committee ships a responsible-AI charter. The data office launches a multi-year readiness program. The strategy offsite debates GPT versus Gemini. Every one of these decisions feels responsible. Every one of them funds an artifact that will not survive contact with how agents actually create or leak value.

That is the pattern, and it is worth naming plainly because it repeats across functions that never talk to each other. In an agentic enterprise, value and risk both live in the workflow: the legible output an agent can read, the outcome you can meter, the permission an agent can exercise, the cross-functional tradeoff an agent quietly makes. The enterprise, meanwhile, keeps measuring and funding a different unit entirely: the internal system, the seat, the principle, the estate, the title. The gap between those two units is where margin leaks, and it compounds quarter after quarter while everyone reports green.

The unit you measure is not the unit that decides

Start with the buyer, because the buyer is what changed. When an AI agent mediates the purchase, product data becomes the interface, and the agent does not read like a person. An audit of product pages from leading brands found that 70% currently do not meet Google’s Universal Commerce Protocol for AI selling. That is the top of the market, functionally invisible at the exact moment the reader at the point of sale became a machine parsing structured inputs.

Here is where the wrong-unit habit shows itself. A senior team hears about the number and comforts itself with the PIM investment. Clean for whom? Internal PIM tidiness and external protocol legibility are different problems, and only one of them earns revenue. A retailer with a sophisticated PIM and a poor protocol-compliant feed gets skipped. A retailer with a basic PIM and a clean feed gets recommended. The unit the catalog team optimized (internal cleanliness) is not the unit the agent grades (protocol-legible output). Pages with structured data are cited 3.1x more frequently in Google AI Overviews, and Shopify reports that orders from AI-powered searches grew 15x year over year through 2025. The demand is arriving whether or not your data is written for the reader who now decides.

Walk down the hall to procurement and the same fracture appears in a different costume. Your largest vendors are changing the thing they sell, moving from a seat license to a charge per resolved ticket, per booked meeting, per completed workflow. Per-seat pricing is structurally broken for AI agents, because the better the agent works, the fewer seats a buyer needs, so the vendor is paid to under-deliver. The market is already validating the move. Sierra reached $100 million in annual recurring revenue in less than two years on an outcome-based approach, and Salesforce priced Agentforce resolutions at two dollars a conversation. Yet procurement still knows how to count users, benchmark a per-user rate, and grind a discount on a term. Ask that team to score a per-outcome contract and the tooling collapses, because the value lives in the workflow while procurement scores the SKU. Consider a single support workflow that once required a help-desk seat, a CRM seat, a knowledge-management seat, a collaboration seat, and a premium copilot entitlement. When an agent absorbs that workflow, you are negotiating five contracts against a piece of work none of them can see.

The oversight slide is the third instance of the same mistake. Every board deck this year has a reassuring box that says human in the loop, and nobody asks who the human is, how many hours they have, and what those hours cost. The gate is trivial to build. The human doing the reviewing is not free and is not infinitely available. A $35-per-hour operations reviewer spending two minutes per approval costs $1.17 per reviewed action. Route a returns agent’s ten thousand flagged actions a day through that reviewer and you have added over $4 million a year in reviewer labor to a workflow you sold to the CFO as pure savings. The enterprise costed the agent and pretended the human was free, because the unit on the business case was the agent, not the workflow of agent plus oversight together.

Why responsible instincts produce the wrong unit

None of this happens because the people involved are careless. It happens because every instinct that felt responsible in the human era now points at the wrong target.

The human era rewarded internal order. A clean PIM, a well-negotiated seat term, a governance charter, a single governed data foundation: these were the marks of a serious operator, and they earned revenue because a human sat at the end of every workflow and could compensate for a rough edge. A shopper who hit vague product copy still browsed. A support rep who held five seats still walked the ticket across all five systems in their head. A reviewer with judgment still caught the odd bad decision. The human was the integration layer, the tradeoff-resolver, and the error-correction routine all at once, and none of that labor appeared on a scorecard because it was assumed.

Agents remove the human from the middle of the workflow, and every assumption that rode on that human comes due at once. The agent reads structured inputs, skips what it cannot parse, and moves on with no notification and no abandoned cart to retarget. The agent runs the whole workflow across all five contracts and exposes that they were priced against a unit of work none of them can see. The agent does not go to the Monday meeting where three VPs used to reconcile their targets. So a marketing personalization agent tuned to conversion, a merchandising pricing agent tuned to margin, and a supply chain replenishment agent tuned to service level each report green while total enterprise margin goes backward. Every agent is locally optimal. The system is globally dumb. Nobody owns the tradeoff between them, because the tradeoff does not live on anyone’s scorecard.

The governance version is the most revealing, because it shows the wrong-unit instinct at its most seductive. An agent is a non-human identity holding credentials, calling tools, and writing to production systems while nobody watches. The real risk surface is what an agent can read, write, and spend. Yet boards fund governance as an ethics project rather than an access-control problem, producing principles documents and a review committee that meets monthly. A SAP LeanIX survey found that 98% of companies have already deployed AI agents or plan to, and less than half of organizations have visibility into an inventory of the agents already running inside them. You cannot govern what you cannot enumerate, and a values statement does not fix an inventory gap. Only plumbing does. Gartner names the failure mode precisely: failures occur when organizations fail to distinguish between an agent’s ability to act and the scope of access it is granted. Capability and permission are two different dials, and most programs turn only one of them.

What the wrong unit costs, and why it compounds

The Cost of Doing Nothing here is a meter running against you, and it runs faster because the enterprise cannot see it.

On the demand side, by 2030 nearly 50% of online shoppers are expected to use AI agents, accounting for roughly 25% of their spending and adding $115B to the US ecommerce sector. Every quarter your catalog stays human-only, a competitor with worse products but cleaner data collects the recommendation. That is margin leaking to a rival’s data hygiene, and there is no abandoned cart to win it back.

On the procurement side, Zylo’s 2026 SaaS Management Index reported that organizations leave an average of 36 percent of SaaS licenses unused against recommended utilization, and now variable outcome charges are being layered on top of that overbought estate. Gartner predicts at least 40 percent of enterprise SaaS spend will shift to usage, agent, or outcome-based models by 2030, with seat-based revenue share declining from 21 percent to 15 percent. Whoever owns the meter owns the invoice, and today the meter, the definition of a valid outcome, and the margin all sit on the vendor’s side of the table.

On the oversight side, the cost is not only the unbudgeted reviewer labor. Reviewers fatigue, attention decays across a queue, and safety becomes an inverted-U in the escalation rate, so past a certain point adding more human oversight makes a system less safe. Gate everything and you can buy worse outcomes while paying full price for oversight you are not actually receiving. As of August 2026, any retail AI system qualifying as high risk, including those controlling consumer credit, BNPL, pricing, or segmentation, must undergo conformity assessments, risk management, and human oversight, with significant penalties for non-compliance. When the incident or the regulator arrives, an unfunded loop becomes the exhibit.

On the governance side, Gartner predicts that by 2027, 40% of enterprises will demote or decommission autonomous AI agents due to governance gaps identified only after production incidents occur. The bill for an over-permissioned agent arrives when it is scaled into a margin-critical workflow and someone finally asks the question your deck never forced. And the readiness program that was supposed to prevent all of this carries its own compounding bill. Gartner projects that through 2026 organizations will abandon 60% of AI projects that lack AI-ready data foundations, and a Cloudera and Harvard Business Review Analytic Services survey of 1,574 IT leaders published in March 2026 found that only 7% say their data is completely ready for AI. Everyone cites those numbers to justify a larger foundation program. Read again: a 60 percent abandonment rate and a 7 percent completion rate are evidence that the big-program model has a catastrophic finish rate, not that you need more of it.

The operating model that fixes it

The fix is a discipline about the unit of account, applied consistently across the functions that currently disagree about it. Four moves, and they reinforce each other.

First, make the workflow the unit of account everywhere. In the catalog, that means a protocol-compliant feed measured on agent citation rate, with complete canonical attributes on your top revenue categories first and real-time price and inventory served through APIs rather than rendered page text. In procurement, it means scoring the workflow and its cost per completed decision rather than the license, and refusing to sign an outcome term you cannot meter yourself, with the Outcome Measurement Agreement as a gating document rather than an appendix. In oversight, it means a second line under every agent labeled reviewer hours, fully loaded, with no line meaning no approval.

Second, scope every readiness effort to a single workload rather than the estate. The teams actually putting agents into production do not try to make the estate ready; they make one workload ready. Pick replenishment, or invoice matching, or the rework loops in the contact center. Assess the specific data sources that one use case requires, fix the schemas and duplicates for that isolated dataset, ship the agent, and let the margin fund the next workload. You will duplicate some effort and accumulate a patchwork a purist will hate. That is the trade: a messier map in exchange for agents in production and a foundation that earns its keep as it grows.

Third, fund plumbing, not posters. Treat every agent as a first-class identity with least-privilege access, just-in-time authorization, and continuous behavioral monitoring. The test is whether you can produce the agent inventory with a named owner for each one, show a specific agent’s exact read, write, and spend scope as a live system query rather than a document, and revoke that agent’s access from one console while someone times it. If revocation takes a ticket, a team, and a Tuesday, your kill switch is aspirational.

Fourth, name the seat that owns the tradeoff. IBM’s 2026 CEO study found that 76% of surveyed organizations now have a Chief AI Officer, up from 26% in 2025, but a title is not accountability, and most of these seats are science-fair curators with a budget and no P&L. Name an Agentic Operating Owner who sits above merchandising, marketing, and supply chain and owns the one thing the functions cannot: the cross-functional tradeoff. Give that seat the P&L for cross-functional agent decisions, the authority to set and revoke autonomy limits per agent per domain in real time, and a kill switch that actually works. Sequence the transition now: the window to reshape vendor terms is closing as this year’s renewals land, and after it you inherit the terms the vendor already wrote.

Everyone reaches roughly the same model frontier, so the retailers who win the next three years will be the ones who decided early and explicitly which unit they are measuring, funding, and holding someone accountable for. Answer one question in a single sentence before your next quarter closes: who owns the tradeoff, and can they turn an agent off in under a minute. If that sentence does not exist yet, the silence is your operating model, and it is already pricing your margin for you.