All repositories

Evals & observabilitycomet-ml

opik

A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.

Who it is for

For teams who want traces and automated evaluations in one place.

How to start

  1. Install and configure: pip install opik, then opik configure.
  2. Wrap a function with the @track decorator so its calls show up as traces.
  3. To self-host, clone https://github.com/comet-ml/opik.git and start the platform per the README.

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+618Sep 6 — Oct 6
21,78122,399

Author's description

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

langfuse

langfuse

An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.

35Kstars+1.2K in 30 dTypeScript

mlflow

mlflow

A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.

28Kstars+455 in 30 dPython

promptfoo

promptfoo

A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.

26Kstars+919 in 30 dTypeScript

openai

evals

OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.

20Kstars+181 in 30 dPython

confident-ai

deepeval

A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.

19Kstars+533 in 30 dPython

vibrantlabsai

ragas

A metrics library for evaluating RAG and other LLM apps: it checks how well an answer relies on retrieved sources. Can also generate test sets.

16Kstars+322 in 30 dPython