An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.
Evals & observabilityUKGovernmentBEIS
inspect_ai
A model-evaluation framework from the UK AI Security Institute: tool use, multi-turn dialog, model-graded scoring and over 200 ready-made evals.
Who it is for
For researchers and engineers who evaluate models and agents seriously.
How to start
- Open the documentation at inspect.aisi.org.uk.
- Browse the catalog of 200+ ready-made evals at inspect.aisi.org.uk/evals.
- For development, clone the repo and run
pip install -e ".[dev]".
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
Inspect: A framework for large language model evaluations
More in «Evals & observability»
A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.
A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.
A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.
OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.
A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.
