An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.
Evals & observabilityfchollet
ARC-AGI
A famous abstract-reasoning test: 800 puzzles on coloured grids where you infer the rule from a few examples. Easy for people, long out of reach for models. The repo holds the data and a page for solving tasks by hand.
Who it is for
For people who want to see how models are tested for real reasoning, not memorised knowledge.
How to start
- Open the
datafolder:traininghas 400 tasks to prototype on,evaluationhas another 400 for the final check. - Open
apps/testing_interface.htmlin a browser and load a task JSON file. - Solve a few tasks yourself, then run your own model on them.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
The Abstraction and Reasoning Corpus
More in «Evals & observability»
A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.
A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.
A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.
OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.
A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.
