Look twice.Find the gem.

AI agents and MCP servers, each published with its source and what the checks found.

Marketplace

  • Everything
  • AI agents
  • Apps
  • MCP servers
  • Templates
  • What people want
  • What changed this week
  • The verification standard
  • The ooruby Index
  • Servers that publish no source
  • Reliability guides
  • What the catalogue holds
  • Sell here

Our library

  • Everything, in one place
  • Guides
  • Glossary
  • Calculators
  • Checklists and cheat sheets
  • Community

ooruby

  • Home
  • For teams
  • Site status
  • Company projects
  • RSS feed

Verification records what our published tests found on a specific version at a specific date. It is not a warranty, and it does not certify that software is free of defects.

AI agentsAppsMCP serversTemplatesWantedCommunityOur library
Sign inSell
Catalogue/Apps/Llmprobe

Llmprobe

ddalcu · App

Llmprobe checks an LLM inference engine from the outside. Given any OpenAI-compatible endpoint, it scores how much of the standard API surface exists, how correct the implemented parts are, and whether the model clears a floor on tool calling, instruction following, JSON output and memory of the conversation. Grading is deterministic, with no model acting as judge. Three multi-step tasks in a simulated file workspace test whether the model can run an agent loop.

Written by ooruby from the README of llmprobe 0.6.16.

Not claimed by its maker. Is this yours? Prove it and answer the findings

Why it's useful

You can tell whether a problem comes from the inference engine or from the model, because coverage, conformance and capability are scored separately.

What you could do with it

  1. 1Point it at a local llama.cpp or LM Studio server to see which API surfaces it implements.
  2. 2Check whether a small model calls tools correctly, and holds back when it should not call one.
  3. 3Run the agent-loop tasks against a model to see whether it can complete multi-step tool work.

Written from its documentation: what it says it can do, not what our checks found.

LLappDeveloper tools

Maintenance

Not measured

No release date is on record for this package yet. The nightly publisher check records one as it reaches each npm package, about once every ten nights.

Release activity only: it says nothing about quality or safety, and a finished small package can be fine without releases. How it is measured

Not yet tested

What this does not cover

  • ·Everything. Treat it as you would software from anywhere else.

Receipt history: every signed receipt for every version of this listing, and what changed between them.

In its maker's words

Conformance and capability test suite for LLM inference engines. Scores surface coverage, spec conformance, and whether the model is semi-capable.

Compare with

  • Node Llama Cppside by side
  • Anthropic Ai Sandbox Runtimeside by side
  • Ruvllmside by side
Compare all 4

How far this has been checked

Static scanned

What it establishes
Its published package, or for a hosted server its source repository at a recorded commit, was read file by file, without running it, by every check in our current scan.
What it does not
How it behaves when it runs, or anything the scan does not read yet: several rubric checks, the repository's history, and compiled code in folders named dist or build. For a hosted server, that the endpoint runs the code that was read. The ooruby Index sets out exactly what was read.
Exactly what was read
Version 0.6.16, the file with digest sha512-ZjSUZTINNh2O1YlEa7tTdymKSJQtVbhGUXpvKN6Jbeir0YwFM3wudREEC2te6CePMDNwdtnUAR1f60PSwqFvjg==, as the registry publishes it. A published version cannot be replaced, so this names the same bytes for anyone who checks.
Authentication
Runs locally. No hosted endpoint is listed for this package. You run it on your own machine, so there is no endpoint of ours to authenticate against.
How the rungs work

Adding it

Run it with:

npx -y llmprobe@0.6.16

Pinned to 0.6.16, which is the version the findings above were found in. Drop the version to take whatever is newest, and the report on this page stops describing what you installed.

Where this came from

Registry
npm
Package
llmprobe
Version read
0.6.16
Resolved
2026-10-10
Open the source(opens in a new tab)

Built from source, attested

npm publishes a signed statement that this exact package was built from the repository below, at this commit. It is verifiable without taking our word for it.

Repository
github.com/ddalcu/llmprobe
Commit
3def676633d9
Built by
github.com/actions/runner/github-hosted
Registry publish attestation
signed by npm · 2026-10-10
Checked
2026-10-10

The badge, if this is your listing

It renders the current rung (static scanned) and links back here, where what that does and does not establish is one click away. It updates itself as the evidence deepens.

[![ooruby: static scanned](https://ooruby.com/api/badge/llmprobe)](https://ooruby.com/market/llmprobe)

About this listing

Kind
App
Category
Developer tools
Pricing
Free
Sandbox
No
Hosted
No
Updated
2026-10-05
Price
Free
Get it freeVisit maker
Rating
no reviews yet
0 want this so far
Sign in to vote

Sign in to watch this listing and hear when its version, rung or findings change.

Similar apps

All apps
CCappDeveloper tools
4 findingsStatic scanned
Besides the terminal, you can use it in your IDE or tag Claude on GitHub. Anthropic collects feedback when you use it, including usage data such as code acceptance, conversation data and bug reports.

Claude Code

anthropic-ai
Free

Puts Claude in your terminal to understand your codebase, edit files and handle git workflows.

80.2M installs/moSEE LICENSE IN README.md
EWappDeveloper tools
2 findingsStatic scanned
Build a collection of providers and stream through it, registering every built-in provider or only the ones you use to keep bundle size down. Tool arguments are parsed as they stream in, and can be checked before a tool runs.

Earendil Works Pi AI

earendil-works
Free

Gives your app one API for streaming from many AI model providers, with typed and validated tools.

14.6M installs/moMIT
TOappDeveloper tools
1 findingStatic scanned
A JavaScript library with no dependencies and a bundled command line tool. It can also check text against a token limit, slice text by token position and split it into chunks.

Tokenx

johannschopplich
Free

Estimates how many tokens a text will use without loading a full tokeniser, from code or the shell.

5.8M installs/moMIT

Verification records what our published tests found on a specific version at a specific date. It is not a warranty, and it does not certify that software is free of defects.