All repositories

Evals & observabilityUKGovernmentBEIS

inspect_ai

A model-evaluation framework from the UK AI Security Institute: tool use, multi-turn dialog, model-graded scoring and over 200 ready-made evals.

Who it is for

For researchers and engineers who evaluate models and agents seriously.

How to start

  1. Open the documentation at inspect.aisi.org.uk.
  2. Browse the catalog of 200+ ready-made evals at inspect.aisi.org.uk/evals.
  3. For development, clone the repo and run pip install -e ".[dev]".

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+227Sep 6 — Oct 5
2,7062,933

Author's description

Inspect: A framework for large language model evaluations

langfuse

langfuse

An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.

35Kstars+1.2K in 30 dTypeScript

mlflow

mlflow

A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.

28Kstars+455 in 30 dPython

promptfoo

promptfoo

A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.

26Kstars+919 in 30 dTypeScript

comet-ml

opik

A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.

22Kstars+618 in 30 dPython

openai

evals

OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.

20Kstars+181 in 30 dPython

confident-ai

deepeval

A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.

19Kstars+533 in 30 dPython