Look twice.Find the gem.

AI agents and MCP servers, each published with its source and what the checks found.

Marketplace

  • Everything
  • AI agents
  • Apps
  • MCP servers
  • Templates
  • What people want
  • What changed this week
  • The verification standard
  • The ooruby Index
  • Servers that publish no source
  • Reliability guides
  • What the catalogue holds
  • Sell here

Our library

  • Everything, in one place
  • Guides
  • Glossary
  • Calculators
  • Checklists and cheat sheets
  • Community

ooruby

  • Home
  • For teams
  • Site status
  • Company projects
  • RSS feed

Verification records what our published tests found on a specific version at a specific date. It is not a warranty, and it does not certify that software is free of defects.

Rubricv1.0
AI agentsAppsMCP serversTemplatesWantedCommunityOur library
Sign inSell
Running it safely
All guides
Running it safelyBeginner· 5 min read

Try a listing before you pay

The one thing to remember

Ten minutes on your own awkward case beats an hour of reading the description.

What you will be able to do

Test a listing on real work of your own and get a defensible answer about whether to buy it.

Figure

How Try a listing before you pay works, in one picture

1Check what can be tried here2Bring the awkward case3Watch what it does, not just what it returns4Run it twice

The same argument as the text, as a chain. Each step is what makes the next one possible.

  1. 1

    Check what can be tried here

    A hosted listing's page, under Connecting to it

    A hosted server can be tested from its page: Test the connection sends the MCP handshake and asks for the tool list, with no credentials and never a tool call, and says whether it answered. An endpoint that asks for a key cannot be tested that way, and its page says so.

    A package cannot be run here yet. Running a package means executing code this site has not finished reading, and that waits on isolated infrastructure the owner has not chosen yet, so no listing carries a sandbox run and none claims one. Run it yourself, in a container you are willing to throw away, and where you cannot, look for the refund window instead.

  2. 2

    Bring the awkward case

    Everything works on a clean, well formed, representative example. The value of a trial run is entirely in the example you would not put in a demo.

    Use the ambiguous ticket, the badly scanned invoice, the meeting where two people talked over each other. That is where the difference between listings actually lives.

  3. 3

    Watch what it does, not just what it returns

    Run it where you can see what it did: which tools it called, what it sent out, how long it took and what it cost. The output tells you whether it is any good. What it did tells you whether you want it near your data.

    A listing that produced a decent answer by calling six tools you did not expect is a different product from one that produced the same answer by calling two.

    If it contacts a destination that is not in the declared egress list, tell us. That is a failed check, and we will look at it again.

  4. 4

    Run it twice

    Agents are not deterministic. One good result is a sample of one, and the variance between runs is often larger than the difference between two listings.

    Two runs on the same input is the cheapest reliability test available, and it takes another two minutes.

You have got it when

You have tried something on your own input, watched what it did, and know what you would do next.

Open the catalogue on ooruby

Read next

Before you buy
How to judge an agent in ten minutes
Running it safely
The first week with a new agent
Running it safely
Work out the blast radius before you grant a scope
Running it safely
Why agents fail in production, and what to watch
The bottom line

Ten minutes on your own awkward case beats an hour of reading the description.

See the AI and semiconductor names

The near-monopolies and the commodities, side by side, because they look identical from outside and they are not.