A skills set and methodology for coding agents: clarify the task, plan, write tests, then code. The agent works with more discipline.
The AI movement's repository database
The best AI repositories
224 living projects — from agent frameworks to prompt-injection defense. Each comes with our own note: why it matters, who it is for and how to start in three steps.
Stars, forks and activity refresh daily via the GitHub API · updated October 6, 2026
- repositories
- 224
- categories
- 11
- stars in total
- 11M
- new stars in 30 days
- +235K
Last 30 days
Growing fastest
The most new stars over the last 30 days.
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
An open coding agent for the terminal with a desktop app. It is not tied to one provider: plug in whichever models you have.
A toolkit for spec-driven development: principles and a spec first, then a plan and tasks, and only then code. Works with several coding agents.
OpenAI's lightweight coding agent that runs in the terminal and works on your project's code. Written in Rust.
Categories
Catalog
14 repositories
An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.
A TypeScript framework for agents and AI apps: agents, workflows, memory and quality evaluation in one place. A good fit for Node-based web teams.
Tooling for computer-use agents: drivers for several OSes, virtual machines and benchmarks for training and evaluation.
A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.
A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.
A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.
Google's code-first agent framework: build agents in Python, run evaluations and deploy. Tuned for Gemini but works with other models too.
A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.
A metrics library for evaluating RAG and other LLM apps: it checks how well an answer relies on retrieved sources. Can also generate test sets.
The standard toolkit for running models through hundreds of academic benchmarks. Many open model leaderboards are built on it.
An observability and evaluation tool for LLMs: starts locally with one command and shows traces, experiments and datasets. Built on OpenTelemetry.
The main benchmark for coding agents: a model must fix real issues from GitHub projects, and the patch is verified with tests in containers.
A model-evaluation framework from the UK AI Security Institute: tool use, multi-turn dialog, model-graded scoring and over 200 ready-made evals.
A benchmark of hard terminal tasks: a set of assignments and a runner where the agent acts as an admin or developer. Has a leaderboard.
Did we miss something?
Suggest a repository
Send a GitHub link and a few words on why it belongs here. Every suggestion is reviewed by hand — not everything gets in.










