An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.
Evals & observabilityAgentOps-AI
agentops
An SDK for observing agents: it records sessions, tracks cost and shows the chain of steps. Integrates with CrewAI, LangChain, OpenAI Agents SDK and others.
Who it is for
For those running agents who want to see what they did and what it cost.
How to start
- Install:
pip install agentops. - Get an API key in the project settings at app.agentops.ai.
- Call
agentops.init(...)with the key at program start and view the session in the dashboard.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Star history is accumulating — the chart appears in a few days.
Author's description
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
More in «Evals & observability»
A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.
A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.
A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.
OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.
A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.
