A skills set and methodology for coding agents: clarify the task, plan, write tests, then code. The agent works with more discipline.
The AI movement's repository database
The best AI repositories
224 living projects — from agent frameworks to prompt-injection defense. Each comes with our own note: why it matters, who it is for and how to start in three steps.
Stars, forks and activity refresh daily via the GitHub API · updated October 6, 2026
- repositories
- 224
- categories
- 11
- stars in total
- 11M
- new stars in 30 days
- +235K
Last 30 days
Growing fastest
The most new stars over the last 30 days.
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
An open coding agent for the terminal with a desktop app. It is not tied to one provider: plug in whichever models you have.
A toolkit for spec-driven development: principles and a spec first, then a plan and tasks, and only then code. Works with several coding agents.
OpenAI's lightweight coding agent that runs in the terminal and works on your project's code. Written in Rust.
Categories
Catalog
9 repositories
A platform where you build agents in a visual builder or by describing them in plain words, then run them on a schedule or a trigger. Hosted version plus free self-hosting.
The core library for working with models: one codebase for text, vision, audio and multimodal tasks, for both inference and training. Most new open models ship through it.
OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.
An IBM library for document parsing: PDF, Office, HTML and images become one unified structure from which Markdown is easy to get. Supports vision models for hard pages.
A local replacement for cloud APIs: an OpenAI-compatible server that runs text, voice and image models on any hardware, no GPU required.
A compact multimodal model that understands images and video and fits on a phone or laptop. Ships an online demo and a web demo you can self-host.
A model that parses a UI screenshot into buttons and fields so an agent can work from the screen alone, without access to page code.
The Qwen team's family of multimodal models: understand images, documents and video, locate objects in a frame and work with UIs.
Meta's Segment Anything 2: picks out any object in an image or video from a click and tracks it across frames. The repo has code, checkpoints and notebooks.
Did we miss something?
Suggest a repository
Send a GitHub link and a few words on why it belongs here. Every suggestion is reviewed by hand — not everything gets in.










