A skills set and methodology for coding agents: clarify the task, plan, write tests, then code. The agent works with more discipline.
The AI movement's repository database
The best AI repositories
224 living projects — from agent frameworks to prompt-injection defense. Each comes with our own note: why it matters, who it is for and how to start in three steps.
Stars, forks and activity refresh daily via the GitHub API · updated October 6, 2026
- repositories
- 224
- categories
- 11
- stars in total
- 11M
- new stars in 30 days
- +235K
Last 30 days
Growing fastest
The most new stars over the last 30 days.
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
An open coding agent for the terminal with a desktop app. It is not tied to one provider: plug in whichever models you have.
A toolkit for spec-driven development: principles and a spec first, then a plan and tasks, and only then code. Works with several coding agents.
OpenAI's lightweight coding agent that runs in the terminal and works on your project's code. Written in Rust.
Categories
Catalog
17 repositories
An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.
A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.
A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.
A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.
OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.
A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.
A metrics library for evaluating RAG and other LLM apps: it checks how well an answer relies on retrieved sources. Can also generate test sets.
The standard toolkit for running models through hundreds of academic benchmarks. Many open model leaderboards are built on it.
An observability and evaluation tool for LLMs: starts locally with one command and shows traces, experiments and datasets. Built on OpenTelemetry.
An OpenTelemetry-based toolkit that automatically traces calls to models, vector databases and frameworks. Data can go to any compatible monitoring system.
A gateway and observability platform: change the API address in your code and all model requests get logged, costed and compared. Can be self-hosted.
The main benchmark for coding agents: a model must fix real issues from GitHub projects, and the patch is verified with tests in containers.
An SDK for observing agents: it records sessions, tracks cost and shows the chain of steps. Integrates with CrewAI, LangChain, OpenAI Agents SDK and others.
A famous abstract-reasoning test: 800 puzzles on coloured grids where you infer the rule from a few examples. Easy for people, long out of reach for models. The repo holds the data and a page for solving tasks by hand.
An observability platform from the makers of Pydantic: it traces plain code, model calls and agents. Built on OpenTelemetry, with an open-source SDK.
A model-evaluation framework from the UK AI Security Institute: tool use, multi-turn dialog, model-graded scoring and over 200 ready-made evals.
A benchmark of hard terminal tasks: a set of assignments and a runner where the agent acts as an admin or developer. Has a leaderboard.
Did we miss something?
Suggest a repository
Send a GitHub link and a few words on why it belongs here. Every suggestion is reviewed by hand — not everything gets in.










