In August 2026 a Fortune 500 company turned on agent discovery for the first time and found 18,000 AI agents active on its endpoints. It had approved 300. That gap became the opening slide of CrowdStrike President Michael Sentonas’ keynote at Fal.Con 2026, and it should be the opening slide of your next engineering review, because the number is what happens when a technology that anyone can install meets an organization that inventories by approval rather than by observation.
Seventeen thousand seven hundred agents sat outside the approved list. That is a ratio of sixty to one. Among the 18,000 that CrowdStrike’s Falcon Guardian identified were Claude Code, OpenAI Codex, Cursor and Kiro, tools your engineers pulled down themselves because the tools made them faster. None of that required a procurement ticket, a security review, or a line in anyone’s system of record. The company was counting the wrong thing.
If you run production agents, or you run the engineers who do, the practical question is how wrong your own count is, and who is responsible for closing the difference before it closes on you. This piece is about that operating problem: how to discover what is actually running, how to inventory it, and how to tier it by what each agent can touch and undo.
The count you trust was never a count
Start with the mechanism that produces a sixty-to-one gap, because it is not exotic. Every enterprise inventory of agents that I have seen is built the same way: it lists what was requested and approved. That works when deployment requires a gate you control, a server to provision, a license to buy, a vendor to onboard. Agents break the assumption. An engineer installs Claude Code on a laptop in the afternoon and it is acting against internal systems by evening. The approval list never learns this happened, so the list and the reality drift apart at the speed of pip install.
The surveys converge on how wide that drift already is. Okta’s AI Agents at Work 2026 research, reported by Yahoo Finance, found that 67% of US workers are currently using unapproved AI tools to do their jobs, and that workers bring their own tools 78% of the time. The visibility mismatch is the part that should worry an operator: 90% of executives express confidence in their visibility into AI tools, while only 11% of those applications are actually visible to IT. Confidence is running roughly eight times ahead of sight. Veeam’s September 2026 survey of EMEA enterprises, covered by Compare the Cloud, puts the same finding from the IT side: 67% of IT departments report that employees are standing up autonomous AI workflows IT cannot track, and 75% of enterprises have no clear oversight of the agents handling their sensitive data.
The trajectory makes the gap a permanent condition rather than a one-time cleanup. Gartner projects that Fortune 500 companies will run more than 150,000 AI agents by 2028, up from fewer than 15 in 2025. CrowdStrike, citing outside research it did not name, puts the current ratio at roughly 90 agents per employee. Whatever the exact multiplier, the direction is set: the population you are trying to govern is growing faster than any approval process staffed by humans can admit new entries. An inventory built on approval will fall further behind every quarter it stays unchanged.
This is shadow IT, and the difference is the verb
Everyone who lived through the shadow-IT decade recognizes the shape. Business units stood up cloud applications without procurement, and IT discovered them later, often during an incident. The Veeam writeup names the parallel directly and then names the twist that matters: where shadow IT leaked documents, shadow agents can rewrite them. That single change in verb is the entire argument for treating agents as a harder problem than the SaaS sprawl that preceded them.
A shadow SaaS app stores. Its worst day is a data exposure: something readable ends up somewhere it should not be, and you contain, notify, and rotate credentials. A shadow agent acts. It reads a ticket, decides, calls an API, moves money, closes an account, edits a record, opens a pull request. The blast radius is no longer the set of documents the tool could see. It is the set of state changes the tool could make, multiplied by the speed at which an autonomous loop makes them. Containment gets harder because there may be nothing sitting still to contain, only a trail of committed actions to reconstruct and reverse.
Three properties compound the risk beyond anything shadow IT presented. The first is identity. Sentonas’ framing is the one to internalize: every agent has an identity, in most cases an overprivileged one, and in too many cases it inherits its human operator’s permissions and gets deployed faster than anyone can govern it. Amazon CISO CJ Moses argued the corollary, that an agent’s scope should never exceed the individual operating it, and that enforcement belongs at the infrastructure layer rather than in the agent’s reasoning. The reason that matters operationally is simple: a human with broad access is one actor making decisions at human speed; an agent inheriting that same access is a tireless actor making decisions in a loop, and only 34% of organizations apply the same security standards to agents as they do to human workers.
The second property is traceability. Per Dataiku and Harris Poll, only 5% of AI output is fully traceable. If an agent acts and you cannot reconstruct why, you cannot audit it, and Cobalt CISO Andrew Obadiaru put his finger on exactly what that costs: when he grants a human access, he is fairly certain what they can do with it and he can audit that access. That certainty is what an untraceable agent removes.
The third property is discovery latency. The median time to detect unauthorized AI tools is 403 days. More than a year passes between an agent going live and anyone in a position to govern it learning it exists. A tool that only stores is dangerous for 403 days in a slow, accumulating way. A tool that acts is making committed changes to production systems for that entire window before it appears on anyone’s radar.
What the gap is already costing
The costs are not hypothetical, and they are not confined to a future breach. ISACA reports that 67% of CISOs dealt with a security incident linked to an unsanctioned AI tool in the last year, and 58% of executives reported a close call. Read those two numbers together: incidents from shadow AI are already the majority experience, not the exception, and the near-misses sit on top of the confirmed hits. Shadow AI costs enterprises an average of $670,000 a year, and when a shadow tool tips into an actual breach the average cost reaches $4.99 million.
There is an operational cost underneath the financial one that engineering leaders feel first. Agents resend their entire working context to the model on every turn. CrowdStrike’s investor briefing puts that traffic at roughly 700 times what a person typing into a browser generates. An unknown agent is therefore an unknown cost center and an unknown load source at the same time. You cannot capacity-plan against a population you cannot see, and you cannot attribute spend to work you did not know was happening. The sixty-to-one gap is a budget line as much as a risk line.
Then there is the posture problem. VentureBeat’s Agentic Security and Identity tracker, July 2026 wave, found that only 18% of 116 enterprises isolate their highest-risk AI agents, and just 8% pair enforcement with isolation. So even among organizations that have started to act, the fraction that both contains its most dangerous agents and enforces limits on what they can do is in single digits. The gap between agents you have approved and agents you are running is compounded by a second gap between agents you can see and agents you have actually constrained.
Discovery first, then everything else
Obadiaru gave the sequence, and it is the right one: the first step is to know what you have, to map and identify the agents, and then map them to their effective permissions. Visibility first, tooling second. An operator should treat any control investment made before discovery as premature, because you cannot enforce policy on a population you have not enumerated.
Discovery has to run where agents actually live, which is the endpoint and the identity layer, because that is where installation and authentication leave traces even when procurement does not. The Fal.Con demonstration illustrated both the promise and the trap: Guardian picked up an engineer’s Claude agent and registered it with the identity provider on its own, so discovery and credentialing happened in the same motion. That is the capability you want, and it is also the moment to be deliberate, because auto-registering an agent as an identity is not the same as authorizing what it can do. Enrollment gives you a name for the thing. It does not yet give you a leash.
Two practices make discovery durable rather than a one-time sweep. The first is treating the discovered inventory as a living feed reconciled continuously against your approved list, so that the difference between the two is a monitored, trending number that a named person watches, not a figure someone stumbles on once a year at day 403. The second is capturing effective permissions at discovery time, not just presence. Knowing that an agent exists tells you almost nothing; knowing that it holds a token that can write to your billing system tells you where to point your attention this week.
There is a demand-side lever that belongs in the same motion, and it is the highest-leverage move available. Research from CSA and Unseen indicates that providing approved AI alternatives can cut unauthorized use by 89%. Discovery tells you which tools your people actually reached for, and 78% of the time they reached for their own. If you sanction fast, good versions of those same tools, the shadow population shrinks by nearly nine tenths on its own, which means discovery is also a product-requirements exercise. The unapproved tools your engineers chose are a ranked list of what to approve next.
Tier by what an agent can touch and what it can undo
Once you can see the population, the governing question is which of your agents can hurt you, and a workable model tiers agents along two axes: what an agent can touch, and whether its actions can be undone.
Touch is the scope of systems and data an agent can reach, and it maps directly to Moses’ rule that an agent’s scope should never exceed its operator’s. An agent that reads a knowledge base and drafts text touches little and reverses cleanly. An agent that holds write access to production databases, payment rails, customer records, or infrastructure sits at the top of the touch axis regardless of how benign its intended task looks in the demo.
Undo is the axis that separates agents from the shadow SaaS that came before, and it is the one most inventories ignore. An action is reversible when a committed change can be rolled back to a known good state without external consequence. Editing a draft is reversible. Sending an email to a customer, executing a payment, deprovisioning an account, deleting records, merging code to a branch that ships: those cross a threshold where the change persists in the world after you notice the mistake. Because a shadow agent can rewrite the documents that shadow IT could only leak, the undo axis is where the real severity lives.
A practical tiering falls out of the two axes. The highest tier is agents that can touch production or sensitive systems and take irreversible actions there; these are the ones that warrant the isolation and enforcement pairing that only 8% of enterprises have achieved, and they should be the first place you spend. The middle tier is agents that touch sensitive systems but only in reversible ways, or that take irreversible actions in low-scope sandboxes; these need monitoring, traceability, and permission review. The lowest tier is read-only, low-scope agents that draft and suggest without committing anything, and these can run under lighter oversight. The point of tiering is to refuse to treat 18,000 agents as one problem. A few hundred of them can end your quarter, and the model’s job is to find those few hundred fast.
Who owns the gap
The reason a sixty-to-one gap can exist is that nobody is measured on closing it. Approval has an owner. Discovery usually does not, and the difference between the two lists sits in the space between security, platform engineering, and the business units doing the installing, which means it belongs to everyone and therefore to no one. That is the failure to fix first, and it is an organizational fix, not a tooling one.
Assign the reconciled gap between approved and discovered agents to a single named owner, with the highest-touch, least-reversible tier as their standing priority. Give that owner the discovery feed, the effective-permissions map, and the authority to constrain or kill an agent at the infrastructure layer where Moses says enforcement belongs. Make the trend line, the gap shrinking or growing week over week, a number that owner reports, because a metric with an owner gets managed and a metric without one becomes a keynote slide at somebody else’s conference.
The governance stack to enforce all of this is arriving on schedule. In a single week in August 2026, Okta Agent SSO, IBM AgentOps and Broadcom AgentMinder reached general availability, covering identity, orchestration, evaluation and runtime action authorization, and CrowdStrike’s Falcon Guardian entered the discovery layer alongside them. The tools to inventory, tier and constrain agents at the infrastructure layer now exist. What has to move ahead of them is the decision to count by observation rather than by approval, and the appointment of the person who owns the distance between the two. Do that in the next planning cycle, before your own discovery sweep returns its first five-figure number, and the gap becomes a managed line item instead of the thing your incident review discovers 403 days too late.