Look twice.Find the gem.

AI agents and MCP servers, each published with its source and what the checks found.

Marketplace

  • Everything
  • AI agents
  • Apps
  • MCP servers
  • Templates
  • What people want
  • What changed this week
  • The verification standard
  • The ooruby Index
  • Servers that publish no source
  • Reliability guides
  • What the catalogue holds
  • Sell here

Our library

  • Everything, in one place
  • Guides
  • Glossary
  • Calculators
  • Checklists and cheat sheets
  • Community

ooruby

  • Home
  • For teams
  • Site status
  • Company projects
  • RSS feed

Verification records what our published tests found on a specific version at a specific date. It is not a warranty, and it does not certify that software is free of defects.

Rubricv1.0
AI agentsAppsMCP serversTemplatesWantedCommunityOur library
Sign inSell
Running it safely
All guides
Running it safelyIntermediate· 7 min read

Why agents fail in production, and what to watch

The one thing to remember

Failures are not spread evenly. They cluster on the inputs you have fewest of, which is why an average success rate tells you almost nothing about your own worst case.

What you will be able to do

Know what to instrument before you rely on an agent, so a bad week shows up as a chart rather than as a complaint.

Figure

How Why agents fail in production works, in one picture

1Understand what a success rate is measuring2Expect the failures to cluster3Watch the retries, because they are the bill4Instrument four numbers, not fourteen5Decide in advance what a bad week looks like

The same argument as the text, as a chain. Each step is what makes the next one possible.

  1. 1

    Understand what a success rate is measuring

    The measured run panel on any listing we have run

    A success rate is the proportion of tasks in a fixed set that produced acceptable output. It is a measurement of that set, on that day, with that model behind it. Independent testing of production agents in 2026 put the average near 57 percent.

    It is still worth more than a marketing claim, because it is a number somebody had to produce by running the thing. Treat it as a floor rather than a forecast.

  2. 2

    Expect the failures to cluster

    Failures are not evenly distributed across your work. They land on the unusual inputs: the long ones, the ones in a second language, the ones with an attachment, the ones from your largest customer who does everything differently.

    Which means an agent at 80 percent overall can be at 20 percent on the cases that matter most, and the average will never tell you. Sample the failures, do not count them.

    If your review queue is a random sample, you will conclude it is fine. Sort by input length and read the tail instead.

  3. 3

    Watch the retries, because they are the bill

    The cost calculator at /tools/ai-agent-cost-calculator

    Every failure that gets retried is billed again. At a 57 percent success rate with two retries, a task costs 1.61 attempts on average, and a metered vendor charges for all of them. Retry until it works and the figure is 1.75.

    The dangerous version is a retry storm: an upstream outage makes every task fail, every failure retries, and the spend for a broken day is higher than for a working one. Cap retries and cap monthly spend on the same afternoon.

  4. 4

    Instrument four numbers, not fourteen

    Attempts per completed task, cost per completed task, the proportion of output a human changed before using it, and time to first useful result. Those four catch nearly everything and fit on one screen.

    The third is the one teams skip and the one that predicts abandonment. An agent whose output is always edited is not saving the time it appears to be saving, and the person doing the editing already knows.

  5. 5

    Decide in advance what a bad week looks like

    Write down the number at which you would turn it off, before you are attached to it. Then check it against the refund window, which on most listings here is fourteen days.

    A decision made in advance is a decision. A decision made in week six is a negotiation with yourself, and the sunk cost always wins that one.

You have got it when

You have four metrics on a dashboard and a written number at which you would stop.

Open the catalogue on ooruby

Read next

Glossary
Success rate
Running it safely
The first week with a new agent
Before you buy
What metered pricing really costs
Glossary
Retry storm
Running it safely
Try a listing before you pay
Running it safely
Work out the blast radius before you grant a scope
The bottom line

Failures are not spread evenly. They cluster on the inputs you have fewest of, which is why an average success rate tells you almost nothing about your own worst case.

See the AI and semiconductor names

The near-monopolies and the commodities, side by side, because they look identical from outside and they are not.