Every vendor blog has a failure taxonomy now. They all list the same things. Timeouts. Retry loops. Infinite loops. Tool errors. Prompt injection. The loud, demo-friendly stuff.
None of that is what kills you in production.
I have deployed enough agents inside real enterprises to know the published lists are written by people who watched agents fail in a sandbox, not by people who got the 11pm call because a customer received a confident, well-formatted, completely wrong answer.
The failures that actually matter do not throw errors. They pass your tests. They look like success. And they cluster into four categories that almost nobody names cleanly.
The four that actually matter
Start with the honest framing. The failure modes are not new to software engineering. Bad inputs, unreliable integrations, cascading errors exist in every production system. What is new is that you have handed routing and execution decisions to a non-deterministic model. The system inherits all the old failure modes plus some new ones, and the model can be confidently wrong in ways that are hard to detect until something is already broken.
Here is the taxonomy that survives contact with production.
Stale data. The agent answers from a snapshot that was true last quarter. The query succeeds. The retrieval succeeds. The number is just old. No error fires because nothing broke, the truth simply moved and the agent did not.
Missing context. This is the big one. Training data gaps compound the problem. Public-internet-trained models lack deep knowledge of internal metrics, lineage tracking, SLAs, and organizational policies. When asked about domain-specific concepts, models fill gaps with plausible but wrong information drawn from statistical patterns. The agent never had your business definition, so it invented one that sounds right.
Hallucination. The classic, but weaponized. When a chatbot hallucinates, a user gets a bad answer and moves on. When an agent hallucinates, it calls delete_user(user_id=4821) on a production database. The fabrication is no longer a paragraph. It is an action with side effects.
Conflict-with-truth. Two systems of record disagree. The CRM says one thing, the ledger says another, and the agent picks one silently and moves on. This is the category no one instruments, and it is the one that produces the most damaging outputs because both answers are internally coherent.
Why your model upgrade makes it worse
This is the part that gets people fired.
The instinct when an agent fails is to reach for a stronger model. That instinct is backwards. Upgrading the model tends to amplify context debt rather than resolve it. A weaker model on wrong context produces obvious errors that are easy to catch in review. A stronger model on the same wrong context produces outputs that are coherent, well-reasoned, and convincingly wrong.
Read that twice. You spend the budget, you ship the smarter model, and you buy yourself failures that are harder to detect. The intelligence you added went straight into making the wrong answer more persuasive.
The root cause is structural, not model-level. AI agents fail in production when the context they run on is assumed rather than governed. That gap, called context debt, surfaces in predictable failure modes: inconsistent answers, authoritative hallucination, tests that pass while downstream reality breaks.
There is a fifth signal that ties it all together. A related indicator is corrections that fail to propagate: the same error category reappears on the next similar query because the correction was never routed back into the context layer. If you fix a wrong answer once and it comes back next week, you do not have a model problem. You have a governance problem.
The Cost of Doing Nothing
Here is where the CODN math bites.
The Cost of Doing Nothing on agent reliability is not a crashed process. A crash is cheap. It is loud, it is logged, it is reproducible, and someone fixes it before lunch.
The real cost is the silent version. AI agents fail silently, completing workflows and returning responses that look correct until downstream consequences reveal the error, often hours later. By the time you find it, the confident wrong answer has been quoted to a customer, baked into a report, and used to make a decision. Then it reappears because the correction never made it home.
So stop optimizing for the demo taxonomy. Build the one that matters. Tag every production failure into stale data, missing context, hallucination, or conflict-with-truth. Route every correction back into a governed context layer, not a prompt tweak. Treat your knowledge base like production code, with ownership and change management.
The teams that win the next two years will not have the smartest models. They will have the cleanest context and the sharpest failure taxonomy. Everyone else will be shipping convincingly wrong answers at scale, and calling it progress.