All repositories

Evals & observabilityopenai

evals

OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.

Who it is for

For those who want to learn how to write and run template-based evals.

How to start

  1. Install the package: pip install evals.
  2. Open docs/run-evals.md and run a ready-made eval.
  3. Read docs/eval-templates.md to write your own.

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+181Sep 6 — Oct 5
19,37319,554

Author's description

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

langfuse

langfuse

An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.

35Kstars+1.2K in 30 dTypeScript

mlflow

mlflow

A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.

28Kstars+455 in 30 dPython

promptfoo

promptfoo

A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.

26Kstars+919 in 30 dTypeScript

comet-ml

opik

A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.

22Kstars+618 in 30 dPython

confident-ai

deepeval

A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.

19Kstars+533 in 30 dPython

vibrantlabsai

ragas

A metrics library for evaluating RAG and other LLM apps: it checks how well an answer relies on retrieved sources. Can also generate test sets.

16Kstars+322 in 30 dPython
openai/evals — what it is and how to start