Look twice.Find the gem.

AI agents and MCP servers, each published with its source and what the checks found.

Marketplace

  • Everything
  • AI agents
  • Apps
  • MCP servers
  • Templates
  • What people want
  • What changed this week
  • The verification standard
  • The ooruby Index
  • Servers that publish no source
  • Reliability guides
  • What the catalogue holds
  • Sell here

Our library

  • Everything, in one place
  • Guides
  • Glossary
  • Calculators
  • Checklists and cheat sheets
  • Community

ooruby

  • Home
  • For teams
  • Site status
  • Company projects
  • RSS feed

Verification records what our published tests found on a specific version at a specific date. It is not a warranty, and it does not certify that software is free of defects.

Rubricv1.0
AI agentsAppsMCP serversTemplatesWantedCommunityOur library
Sign inSell
Back to the standard
rubric v1.0

Measured run

Run against a fixed task set. Success rate, median time and cost per run recorded, never estimated.

Applies to
agents
Standard references
none mapped

A finding here is not a failure

Software of any size carries something. A listing that raises findings on this check is still sold, with the findings published and their severity and reachability stated. A badge that only ever said pass would teach people to stop reading it.

Explainer · 2026-09-30

Measured runs: why an agent's success rate has to come from runs, not demos

An agent chooses its own steps, so the only way to know how often it succeeds is to run it, many times, on a fixed set of tasks, and count.

What this row checks

That the agent was run against a fixed task set, with the success rate, the median time and the cost per run recorded from those runs and never estimated.

Why a demo is not evidence

A demo is one sample from a distribution, usually the best one. Two runs of the same agent on the same task can take different steps and reach different results, and the difference between runs is often larger than the difference between two products. A success rate over a fixed task set is the smallest honest summary of that distribution.

The median is recorded rather than the mean for time because a handful of slow runs would otherwise hide the typical one, and cost is recorded per run because that is the number a buyer multiplies.

What a catalogue listing can show

Nothing, and it says so. A listing indexed from a public registry has never been run here; this row is answerable only for a maker-submitted listing run in the sandbox. A figure that was not measured is shown as not measured, never as zero.

Sources

  1. ooruby: the verification standard
  2. ooruby: what each kind of evidence can answer

In the glossary: Sandbox run, Success rate.

What this check has found

Nothing in the catalogue has raised a finding on this check. That is a fact about what has been tested so far rather than a guarantee about what is out there, and it is printed because a check that never fires is worth knowing about too.

Other checks

No shipped credentialsDependency advisoriesBehaviour matches the manifestDeclared egressData handling disclosedLicence and provenance

Every listing in the catalogue shows its result on this check, with the findings summarised in public and the full report to whoever bought it. Open the catalogue.