Puts Claude in your terminal to understand your codebase, edit files and handle git workflows.
Deepeval
penguine-ip · App
DeepEval for TypeScript is an open-source framework for testing LLM applications such as agents, RAG pipelines and chatbots. It runs as Vitest, so you write test cases and assert that they pass metrics for agent tasks, retrieval, multi-turn conversations, MCP use, images, bias, toxicity and PII leakage. Metrics run on your machine with an LLM as judge, and results can optionally go to Confident AI to compare runs over time.
Written by ooruby from the README of deepeval 0.9.22.
Not claimed by its maker. Is this yours? Prove it and answer the findings
Why it's useful
You can test an LLM app the way you test ordinary code, with checks that pass or fail and explain each failure.
What you could do with it
- Write a Vitest test that fails when your chatbot's answer scores below your relevancy threshold.
- Wrap your agent's functions for tracing and score its full trajectory with the task completion metric.
- Describe your own pass criteria in plain English with G-Eval and score a batch of test cases from a script.
Written from its documentation: what it says it can do, not what our checks found.
Maintenance
Not measuredNo release date is on record for this package yet. The nightly publisher check records one as it reaches each npm package, about once every ten nights.
Release activity only: it says nothing about quality or safety, and a finished small package can be fine without releases. How it is measured
- Licence and provenance
- finding
- no LICENCE file at the root (.)
- Data handling disclosed
- finding
- no statement of what it reads, writes or sends (README.md)
Published because it was found. A finding is information about where the limits are, not a verdict that the software is unsafe.
What this does not cover
- Everything. Treat it as you would software from anywhere else.
Receipt history: every signed receipt for every version of this listing, and what changed between them.
In its maker's words
The LLM Evaluation Framework for TypeScript
Compare with
Compare all 4How far this has been checked
Static scanned
- What it establishes
- Its published package, or for a hosted server its source repository at a recorded commit, was read file by file, without running it, by every check in our current scan.
- What it does not
- How it behaves when it runs, or anything the scan does not read yet: several rubric checks, the repository's history, and compiled code in folders named dist or build. For a hosted server, that the endpoint runs the code that was read. The ooruby Index sets out exactly what was read.
- Exactly what was read
- Version 0.9.22, the file with digest sha512-5JJf1cFsVkBmXYIRU+IRh0Tz5GEKmq6MtTzuzvpHmFbyA4+LckWTcWnZfxCLuffO2RJdsb2bIB0I321LdmzXOg==, as the registry publishes it. A published version cannot be replaced, so this names the same bytes for anyone who checks.
- Authentication
- Runs locally. No hosted endpoint is listed for this package. You run it on your own machine, so there is no endpoint of ours to authenticate against.
Adding it
Run it with:
npx -y deepeval@0.9.22Pinned to 0.9.22, which is the version the findings above were found in. Drop the version to take whatever is newest, and the report on this page stops describing what you installed.
Where this came from
- Registry
- npm
- Package
- deepeval
- Version read
- 0.9.22
- Resolved
- 2026-10-09
No build provenance published. The repository above is the one the publisher declared, and nothing links it to the package you would install. That is not a mark against this listing, since most packages are published this way, but it is a check nobody can run.
The badge, if this is your listing
It renders the current rung (static scanned) and links back here, where what that does and does not establish is one click away. It updates itself as the evidence deepens.
[](https://ooruby.com/market/deepeval)About this listing
- Kind
- App
- Category
- Developer tools
- Pricing
- Free
- Sandbox
- No
- Hosted
- No
- Updated
- 2026-10-08
Similar apps
All appsGives your app one API for streaming from many AI model providers, with typed and validated tools.
Estimates how many tokens a text will use without loading a full tokeniser, from code or the shell.
Verification records what our published tests found on a specific version at a specific date. It is not a warranty, and it does not certify that software is free of defects.