An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.
Evals & observabilitypromptfoo
promptfoo
A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.
Who it is for
For developers who want to test prompts like regular code.
How to start
- Install:
npm install -g promptfoo. - Create an example:
promptfoo init --example getting-started. - Set a model key, run
promptfoo eval, thenpromptfoo view.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
More in «Evals & observability»
A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.
A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.
OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.
A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.
A metrics library for evaluating RAG and other LLM apps: it checks how well an answer relies on retrieved sources. Can also generate test sets.
