Look twice.Find the gem.

AI agents and MCP servers, each published with its source and what the checks found.

Marketplace

  • Everything
  • AI agents
  • Apps
  • MCP servers
  • Templates
  • What people want
  • What changed this week
  • The verification standard
  • The ooruby Index
  • Servers that publish no source
  • Reliability guides
  • What the catalogue holds
  • Sell here

Our library

  • Everything, in one place
  • Guides
  • Glossary
  • Calculators
  • Checklists and cheat sheets
  • Community

ooruby

  • Home
  • For teams
  • Site status
  • Company projects
  • RSS feed

Verification records what our published tests found on a specific version at a specific date. It is not a warranty, and it does not certify that software is free of defects.

AI agentsAppsMCP serversTemplatesWantedCommunityOur library
Sign inSell
Catalogue/Apps/Dt Evals

Dt Evals

dynatrace-oss · App

dt-evals runs continuous evaluations of LLM applications with production traces as the dataset, for teams shipping chat, RAG and agent workflows. It pulls generative AI spans from your Dynatrace environment, samples them, masks personal data, scores each interaction with an LLM judge using 15 built-in or your own metrics, and writes the scores back to Dynatrace, where low scores can be traced to their prompts, retrieval and tool calls. A TypeScript library runs the same evaluators in code.

Written by ooruby from the README of @dynatrace-oss/dt-evals 0.3.3-alpha.

Not claimed by its maker. Is this yours? Prove it and answer the findings

Why it's useful

You can evaluate an LLM app on the traffic it actually handles and trace a falling score to its cause where you already monitor it.

What you could do with it

  1. 1Score the last hour of production chat traces for relevance and faithfulness, then open low scorers in Dynatrace.
  2. 2Fail a CI job when evaluation scores breach your thresholds, using the JSON output and exit codes.
  3. 3Write a custom evaluator from your own prompt and test it from the command line before scheduling runs.

Written from its documentation: what it says it can do, not what our checks found.

DEappInfrastructure

Maintenance

Not measured

No release date is on record for this package yet. The nightly publisher check records one as it reaches each npm package, about once every ten nights.

Release activity only: it says nothing about quality or safety, and a finished small package can be fine without releases. How it is measured

Not yet tested
Licence and provenance
finding
no LICENCE file at the root (.)

Published because it was found. A finding is information about where the limits are, not a verdict that the software is unsafe.

What this does not cover

  • ·Everything. Treat it as you would software from anywhere else.

Receipt history: every signed receipt for every version of this listing, and what changed between them.

In its maker's words

Evaluation CLI for AI Observability on Dynatrace

Compare with

  • Omnirouteside by side
  • AI Ctrfside by side
  • Gatewayside by side
Compare all 4

How far this has been checked

Static scanned

What it establishes
Its published package, or for a hosted server its source repository at a recorded commit, was read file by file, without running it, by every check in our current scan.
What it does not
How it behaves when it runs, or anything the scan does not read yet: several rubric checks, the repository's history, and compiled code in folders named dist or build. For a hosted server, that the endpoint runs the code that was read. The ooruby Index sets out exactly what was read.
Exactly what was read
Version 0.3.3-alpha, the file with digest sha512-VRKFBuomrC2qgCdS+Keyzf5GSaPoWrbePXvRjO6f3ZNXhbgVbXxYKgj4gV26oVRj8IFpkpEHY+v3oJZ6/Lk+9g==, as the registry publishes it. A published version cannot be replaced, so this names the same bytes for anyone who checks.
Authentication
Runs locally. No hosted endpoint is listed for this package. You run it on your own machine, so there is no endpoint of ours to authenticate against.
How the rungs work

Adding it

Run it with:

npx -y @dynatrace-oss/dt-evals@0.3.3-alpha

Pinned to 0.3.3-alpha, which is the version the findings above were found in. Drop the version to take whatever is newest, and the report on this page stops describing what you installed.

Where this came from

Registry
npm
Package
@dynatrace-oss/dt-evals
Version read
0.3.3-alpha
Resolved
2026-10-09
Open the source(opens in a new tab)

Built from source, attested

npm publishes a signed statement that this exact package was built from the repository below, at this commit. It is verifiable without taking our word for it.

Repository
github.com/dynatrace-oss/dt-evals
Commit
e1c029ad42b8
Built by
github.com/actions/runner/github-hosted
Registry publish attestation
signed by npm · 2026-10-10
Checked
2026-10-10

The badge, if this is your listing

It renders the current rung (static scanned) and links back here, where what that does and does not establish is one click away. It updates itself as the evidence deepens.

[![ooruby: static scanned](https://ooruby.com/api/badge/dynatrace-oss-dt-evals)](https://ooruby.com/market/dynatrace-oss-dt-evals)

About this listing

Kind
App
Category
Infrastructure
Pricing
Free
Sandbox
No
Hosted
No
Updated
2026-10-08
Price
Free
Get it freeVisit maker
Rating
no reviews yet
0 want this so far
Sign in to vote

Sign in to watch this listing and hear when its version, rung or findings change.

More from dynatrace-oss

See the storefront
DYMCP serverInfrastructure
Clean scanStatic scanned
Deprecated by its maker, who points to Dynatrace-for-AI with dtctl and the Dynatrace Remote MCP Server instead. Runs on your machine and signs in through your browser; Grail queries can incur costs under your licence.

Dynatrace MCP Server

dynatrace-oss
Free

Lets your agent list Dynatrace problems, vulnerabilities and exceptions and run DQL queries.

No reviews yetMIT

Similar apps

All apps
OSappInfrastructure
Clean scanStatic scanned
Serves a local web dashboard that refreshes every 30 seconds and keeps its figures in a SQLite database. Only new or changed log entries are processed, and results can be broken down by model or by folder.

Omp Stats

oh-my-pi
Free

Reads your local AI session logs and charts requests, errors, tokens per second and cache rate.

352.5k installs/moMIT
OMappInfrastructure
Clean scanPublisher verified
Heavy chat requests queue instead of failing, a shared quota is split fairly across pooled keys, and long context can be compressed. Cost and usage headers come back from every endpoint, with per-key dollar spend quotas.

Omniroute

diegosouza.pw
Free

Routes your AI requests across many providers through one OpenAI-compatible API.

316.3k installs/moMIT
ACappInfrastructure
Clean scanStatic scanned
Works with OpenAI, Anthropic, Gemini, Mistral, DeepSeek, OpenRouter and OpenAI-compatible services such as Ollama, using your own API keys. It can also sum up the whole run.

AI Ctrf

ctrf-io
Free

Adds an AI explanation of each failed test to your CTRF test report, using the model provider you pick.

48.5k installs/moMIT

Verification records what our published tests found on a specific version at a specific date. It is not a warranty, and it does not certify that software is free of defects.