A skills set and methodology for coding agents: clarify the task, plan, write tests, then code. The agent works with more discipline.
The AI movement's repository database
The best AI repositories
224 living projects — from agent frameworks to prompt-injection defense. Each comes with our own note: why it matters, who it is for and how to start in three steps.
Stars, forks and activity refresh daily via the GitHub API · updated October 6, 2026
- repositories
- 224
- categories
- 11
- stars in total
- 11M
- new stars in 30 days
- +235K
Last 30 days
Growing fastest
The most new stars over the last 30 days.
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
An open coding agent for the terminal with a desktop app. It is not tied to one provider: plug in whichever models you have.
A toolkit for spec-driven development: principles and a spec first, then a plan and tasks, and only then code. Works with several coding agents.
OpenAI's lightweight coding agent that runs in the terminal and works on your project's code. Written in Rust.
Categories
Catalog
14 repositories
Repository of the large open DeepSeek-V3 mixture-of-experts model: description, benchmark results and run instructions. It showed an open model can stand next to closed ones.
A free Hugging Face course on agents in four units: fundamentals, the smolagents, LlamaIndex and LangGraph frameworks, agentic RAG and a final benchmark assignment.
Tooling for computer-use agents: drivers for several OSes, virtual machines and benchmarks for training and evaluation.
A research agent from Princeton and Stanford: it takes a GitHub issue and tries to fix it. It set the bar on SWE-bench; the authors now recommend mini-swe-agent.
OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.
The standard toolkit for running models through hundreds of academic benchmarks. Many open model leaderboards are built on it.
The main benchmark for coding agents: a model must fix real issues from GitHub projects, and the patch is verified with tests in containers.
An SDK for observing agents: it records sessions, tracks cost and shows the chain of steps. Integrates with CrewAI, LangChain, OpenAI Agents SDK and others.
A famous abstract-reasoning test: 800 puzzles on coloured grids where you infer the rule from a few examples. Easy for people, long out of reach for models. The repo holds the data and a page for solving tasks by hand.
Meta's generative AI safety project: the Llama Guard and Prompt Guard filter models, the Code Shield scanner and the CyberSec Eval test suites. It pairs defense with attack-based testing.
A suite of eight models of different sizes trained on identical data, with 154 saved checkpoints each. It exists to study how a model learns and when knowledge appears in it.
A model-evaluation framework from the UK AI Security Institute: tool use, multi-turn dialog, model-graded scoring and over 200 ready-made evals.
A benchmark of hard terminal tasks: a set of assignments and a runner where the agent acts as an admin or developer. Has a leaderboard.
A dynamic environment for evaluating attacks and defenses against LLM agents. It runs an agent on user tasks and checks whether a malicious instruction gets through.
Did we miss something?
Suggest a repository
Send a GitHub link and a few words on why it belongs here. Every suggestion is reviewed by hand — not everything gets in.










