An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.
Evals & observabilityopenai
evals
OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.
Who it is for
For those who want to learn how to write and run template-based evals.
How to start
- Install the package:
pip install evals. - Open
docs/run-evals.mdand run a ready-made eval. - Read
docs/eval-templates.mdto write your own.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
More in «Evals & observability»
A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.
A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.
A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.
A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.
A metrics library for evaluating RAG and other LLM apps: it checks how well an answer relies on retrieved sources. Can also generate test sets.
