Why agents fail in production, and what to watch
Failures are not spread evenly. They cluster on the inputs you have fewest of, which is why an average success rate tells you almost nothing about your own worst case.
Know what to instrument before you rely on an agent, so a bad week shows up as a chart rather than as a complaint.
How Why agents fail in production works, in one picture
The same argument as the text, as a chain. Each step is what makes the next one possible.
- 1
Understand what a success rate is measuring
The measured run panel on any listing we have runA success rate is the proportion of tasks in a fixed set that produced acceptable output. It is a measurement of that set, on that day, with that model behind it. Independent testing of production agents in 2026 put the average near 57 percent.
It is still worth more than a marketing claim, because it is a number somebody had to produce by running the thing. Treat it as a floor rather than a forecast.
- 2
Expect the failures to cluster
Failures are not evenly distributed across your work. They land on the unusual inputs: the long ones, the ones in a second language, the ones with an attachment, the ones from your largest customer who does everything differently.
Which means an agent at 80 percent overall can be at 20 percent on the cases that matter most, and the average will never tell you. Sample the failures, do not count them.
If your review queue is a random sample, you will conclude it is fine. Sort by input length and read the tail instead.
- 3
Watch the retries, because they are the bill
The cost calculator at /tools/ai-agent-cost-calculatorEvery failure that gets retried is billed again. At a 57 percent success rate with two retries, a task costs 1.61 attempts on average, and a metered vendor charges for all of them. Retry until it works and the figure is 1.75.
The dangerous version is a retry storm: an upstream outage makes every task fail, every failure retries, and the spend for a broken day is higher than for a working one. Cap retries and cap monthly spend on the same afternoon.
- 4
Instrument four numbers, not fourteen
Attempts per completed task, cost per completed task, the proportion of output a human changed before using it, and time to first useful result. Those four catch nearly everything and fit on one screen.
The third is the one teams skip and the one that predicts abandonment. An agent whose output is always edited is not saving the time it appears to be saving, and the person doing the editing already knows.
- 5
Decide in advance what a bad week looks like
Write down the number at which you would turn it off, before you are attached to it. Then check it against the refund window, which on most listings here is fourteen days.
A decision made in advance is a decision. A decision made in week six is a negotiation with yourself, and the sunk cost always wins that one.
You have four metrics on a dashboard and a written number at which you would stop.
Read next
Failures are not spread evenly. They cluster on the inputs you have fewest of, which is why an average success rate tells you almost nothing about your own worst case.
The near-monopolies and the commodities, side by side, because they look identical from outside and they are not.