A skills set and methodology for coding agents: clarify the task, plan, write tests, then code. The agent works with more discipline.
The AI movement's repository database
The best AI repositories
224 living projects — from agent frameworks to prompt-injection defense. Each comes with our own note: why it matters, who it is for and how to start in three steps.
Stars, forks and activity refresh daily via the GitHub API · updated October 6, 2026
- repositories
- 224
- categories
- 11
- stars in total
- 11M
- new stars in 30 days
- +235K
Last 30 days
Growing fastest
The most new stars over the last 30 days.
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
An open coding agent for the terminal with a desktop app. It is not tied to one provider: plug in whichever models you have.
A toolkit for spec-driven development: principles and a spec first, then a plan and tasks, and only then code. Works with several coding agents.
OpenAI's lightweight coding agent that runs in the terminal and works on your project's code. Written in Rust.
Categories
Catalog
19 repositories
The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.
A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.
OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.
Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.
Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.
A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.
Instant voice cloning: takes a short sample, transfers the timbre to any text and lets you control style, emotion and language. MIT-licensed and free for commercial use.
The standard PyTorch library for diffusion models: ready-made pipelines for images, video and audio that run in a few lines.
An open text-to-speech model with multilingual synthesis and voice cloning. Comes with docs, a Docker image and a server mode.
Resemble AI's open text-to-speech model with voice cloning from a short sample. One command from PyPI to install.
A compact multimodal model that understands images and video and fits on a phone or laptop. Ships an online demo and a web demo you can self-host.
The official inference code for FLUX.1 models: text-to-image generation and editing, including Kontext mode where you edit an image with words.
Whisper on the CTranslate2 engine: same output, noticeably faster and lighter on memory, including 8-bit mode.
The Qwen team's family of multimodal models: understand images, documents and video, locate objects in a frame and work with UIs.
Meta's Segment Anything 2: picks out any object in an image or video from a click and tracks it across frames. The repo has code, checkpoints and notebooks.
Open video-generation models: turn text or an image into a clip. The 1.3B version fits in roughly 8 GB of VRAM.
A Python framework for real-time voice and multimodal agents: compose a speech-model-speech pipeline from modules. Client SDKs for web and mobile included.
LiveKit's framework for realtime voice AI agents: the agent joins a call or room, listens, thinks and answers by voice.
A lightweight 82M-parameter text-to-speech model that sounds far better than its size suggests and runs fast locally.
Did we miss something?
Suggest a repository
Send a GitHub link and a few words on why it belongs here. Every suggestion is reviewed by hand — not everything gets in.










