All repositories

Evals & observabilitylangfuse

langfuse

An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.

Who it is for

For teams who want to understand what their LLM app does in production.

How to start

  1. Create a project in Langfuse Cloud or self-host: git clone --depth=1 https://github.com/langfuse/langfuse.git, then docker compose up.
  2. Install the SDK: pip install langfuse openai and set LANGFUSE_SECRET_KEY and LANGFUSE_PUBLIC_KEY in .env.
  3. Make your first model call and open the trace in the UI.

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+1,152Sep 7 — Oct 5
34,23435,386

Author's description

🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.

mlflow

mlflow

A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.

28Kstars+443 in 30 dPython

promptfoo

promptfoo

A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.

26Kstars+876 in 30 dTypeScript

comet-ml

opik

A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.

22Kstars+593 in 30 dPython

openai

evals

OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.

20Kstars+175 in 30 dPython

confident-ai

deepeval

A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.

19Kstars+506 in 30 dPython

vibrantlabsai

ragas

A metrics library for evaluating RAG and other LLM apps: it checks how well an answer relies on retrieved sources. Can also generate test sets.

16Kstars+306 in 30 dPython