A skills set and methodology for coding agents: clarify the task, plan, write tests, then code. The agent works with more discipline.
The AI movement's repository database
The best AI repositories
224 living projects — from agent frameworks to prompt-injection defense. Each comes with our own note: why it matters, who it is for and how to start in three steps.
Stars, forks and activity refresh daily via the GitHub API · updated October 6, 2026
- repositories
- 224
- categories
- 11
- stars in total
- 11M
- new stars in 30 days
- +235K
Last 30 days
Growing fastest
The most new stars over the last 30 days.
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
An open coding agent for the terminal with a desktop app. It is not tied to one provider: plug in whichever models you have.
A toolkit for spec-driven development: principles and a spec first, then a plan and tasks, and only then code. Works with several coding agents.
OpenAI's lightweight coding agent that runs in the terminal and works on your project's code. Written in Rust.
Categories
Catalog
38 repositories
The easiest way to run an open model on your own machine: one command downloads and starts it, plus a local REST API and libraries for Python and JavaScript.
The core library for working with models: one codebase for text, vision, audio and multimodal tasks, for both inference and training. Most new open models ship through it.
A self-hosted ChatGPT-style interface that connects to Ollama and any OpenAI-compatible API. Installs with a single Docker command and runs on your own server.
A C/C++ engine that runs language models on ordinary hardware: laptops, phones, GPU-less servers. Much of local AI, Ollama included, is built on it.
Repository of the large open DeepSeek-V3 mixture-of-experts model: description, benchmark results and run instructions. It showed an open model can stand next to closed ones.
The foundation under most modern neural networks: GPU-accelerated tensors and automatic differentiation, wrapped in ordinary Python. Transformers, Llama and nearly every open model are written on it.
An engine for fast model serving on GPUs: frugal with memory and handles many requests at once. The default pick when a model has to run as a service.
The open DeepSeek-R1 model, trained with reinforcement learning to reason step by step, plus its smaller distilled versions on Qwen and Llama.
A tool for running and fine-tuning models while saving memory: there is a desktop app, a web UI and a library. Fine-tuning works on a modest GPU.
A unified framework for fine-tuning 100+ models with no code: through a config-driven CLI or the LLaMA Board web UI. Supports LoRA and other methods.
A high-level deep learning library that runs on top of JAX, TensorFlow or PyTorch: write a model once and pick the backend that suits the job.
A minimal, readable GPT implementation in PyTorch by Andrej Karpathy: training from scratch and fine-tuning in a few hundred lines. The best way to see how a language model works.
A single gateway to 100+ LLM providers in OpenAI format: a Python SDK and a proxy server with cost tracking, load balancing and logging. Swap models without changing code.
The whole path from zero to your own ChatGPT-like chat in one repo: tokenizer, pretraining, fine-tuning and UI. Designed to train for roughly a hundred dollars.
A local replacement for cloud APIs: an OpenAI-compatible server that runs text, voice and image models on any hardware, no GPU required.
Joins several devices into one AI cluster so you can run models too big for a single machine's memory. Devices discover each other automatically; there is a dashboard and an API.
A desktop ChatGPT alternative that works fully offline: download a model and chat, with nothing leaving your machine.
An open platform for training, serving and evaluating chat models. It produced Vicuna and runs Chatbot Arena, where people blind-compare model answers.
A fast engine for serving language and multimodal models. A vLLM rival, especially strong on large models like DeepSeek.
GPT-2 training in plain C and CUDA, no PyTorch. Shows what happens under the hood while a model learns.
Apple's array framework for its own chips: a NumPy- and PyTorch-like interface that uses the Mac's unified memory. The base for running and training models on a Mac.
A layer on top of PyTorch that reduces training a network to a few lines. It comes with the well-known free fast.ai course, which makes it a handy place to start.
The open Qwen3 model family from Alibaba, from small to large, with reasoning modes. The README collects run examples for Transformers, llama.cpp and Ollama.
Packs a model and its engine into one executable file: download, make it executable, run. Works across operating systems with no install.
A fast, memory-efficient implementation of attention that large-model training and inference on GPUs lean on. The result is exact, with no approximation.
A library for working with datasets: one line loads thousands of ready-made sets from the Hub or your local files and processes them fast. It can stream data without a full download.
A parameter-efficient fine-tuning library: LoRA and related methods train only a small fraction of the weights. Large models fine-tune on consumer hardware.
Two open OpenAI models, gpt-oss-120b and gpt-oss-20b, plus reference run implementations including Metal for Mac. A rare case of OpenAI releasing weights.
Hugging Face's library for post-training models: SFT, DPO, GRPO and reward-model training. The base for shaping model behavior after pretraining.
An alternative to the transformer: a state space model whose runtime grows linearly with text length. A good fit for long sequences.
Did we miss something?
Suggest a repository
Send a GitHub link and a few words on why it belongs here. Every suggestion is reviewed by hand — not everything gets in.










