Puts Claude in your terminal to understand your codebase, edit files and handle git workflows.
Pipelinescore CLI
pipelinescore · App
PipelineScore CLI runs a fixed benchmark of language models against a model you serve locally or a hosted model reached with your own API key, then submits the result to a public leaderboard grouped by model and hardware. Tasks cover code, reasoning, tool use, grounded answers and speed, with answers checked by programs instead of a judge model. The newer suite has 71 tasks including repository bug fixes and multi-step agent work, and model-written code runs in a Docker sandbox with no network, so Docker must be running.
Written by ooruby from the README of @pipelinescore/cli 0.5.0.
Not claimed by its maker. Is this yours? Prove it and answer the findings
Why it's useful
It scores a model on the hardware you actually use, with answers graded by programs, so runs of the same model on the same hardware can be compared.
What you could do with it
- Point it at a model served by Ollama or LM Studio and score it on the machine you actually run it on.
- Run the benchmark without publishing, so the results stay on your own machine.
- Benchmark a hosted Anthropic or OpenAI model by passing your own API key.
Written from its documentation: what it says it can do, not what our checks found.
Maintenance
Not measuredNo release date is on record for this package yet. The nightly publisher check records one as it reaches each npm package, about once every ten nights.
Release activity only: it says nothing about quality or safety, and a finished small package can be fine without releases. How it is measured
What this does not cover
- Everything. Treat it as you would software from anywhere else.
Receipt history: every signed receipt for every version of this listing, and what changed between them.
In its maker's words
PipelineScore CLI — benchmark LLMs on your own hardware. Zero-config: `npx @pipelinescore/cli` auto-detects Ollama / LM Studio / llama.cpp / MLX and walks you through a run. Fully deterministic suite (no API key required), scored locally, published to the
Compare with
Compare all 4How far this has been checked
Static scanned
- What it establishes
- Its published package, or for a hosted server its source repository at a recorded commit, was read file by file, without running it, by every check in our current scan.
- What it does not
- How it behaves when it runs, or anything the scan does not read yet: several rubric checks, the repository's history, and compiled code in folders named dist or build. For a hosted server, that the endpoint runs the code that was read. The ooruby Index sets out exactly what was read.
- Exactly what was read
- Version 0.5.0, the file with digest sha512-7tpidTvyxamVNnZQo6tbNgQHYSxQLrIOAZcTgcENugFMBBCM7Jl35q3uPyJWyMQmbbUJd6ysyZFlVYpZGeQ1iA==, as the registry publishes it. A published version cannot be replaced, so this names the same bytes for anyone who checks.
- Authentication
- Runs locally. No hosted endpoint is listed for this package. You run it on your own machine, so there is no endpoint of ours to authenticate against.
Adding it
Run it with:
npx -y -p @pipelinescore/cli@0.5.0 ps-benchPinned to 0.5.0, which is the version the findings above were found in. Drop the version to take whatever is newest, and the report on this page stops describing what you installed.
Where this came from
- Registry
- npm
- Package
- @pipelinescore/cli
- Version read
- 0.5.0
- Resolved
- 2026-10-10
No build provenance published. The repository above is the one the publisher declared, and nothing links it to the package you would install. That is not a mark against this listing, since most packages are published this way, but it is a check nobody can run.
The badge, if this is your listing
It renders the current rung (static scanned) and links back here, where what that does and does not establish is one click away. It updates itself as the evidence deepens.
[](https://ooruby.com/market/pipelinescore-cli)About this listing
- Kind
- App
- Category
- Developer tools
- Pricing
- Free
- Sandbox
- No
- Hosted
- No
- Updated
- 2026-10-08
Similar apps
All appsGives your app one API for streaming from many AI model providers, with typed and validated tools.
Estimates how many tokens a text will use without loading a full tokeniser, from code or the shell.
Verification records what our published tests found on a specific version at a specific date. It is not a warranty, and it does not certify that software is free of defects.