A skills set and methodology for coding agents: clarify the task, plan, write tests, then code. The agent works with more discipline.
The AI movement's repository database
The best AI repositories
224 living projects — from agent frameworks to prompt-injection defense. Each comes with our own note: why it matters, who it is for and how to start in three steps.
Stars, forks and activity refresh daily via the GitHub API · updated October 6, 2026
- repositories
- 224
- categories
- 11
- stars in total
- 11M
- new stars in 30 days
- +235K
Last 30 days
Growing fastest
The most new stars over the last 30 days.
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
An open coding agent for the terminal with a desktop app. It is not tied to one provider: plug in whichever models you have.
A toolkit for spec-driven development: principles and a spec first, then a plan and tasks, and only then code. Works with several coding agents.
OpenAI's lightweight coding agent that runs in the terminal and works on your project's code. Written in Rust.
Categories
Catalog
19 repositories
A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A ready-made RAG engine with a UI: it parses documents with templates, chunks them, searches with citations and supports agentic retrieval. Deploys with Docker.
An open-source Python crawler: it visits pages with a real browser and returns clean Markdown for LLMs. Runs locally, in Docker or as a cloud service.
Parses complex PDFs, scans and Office files into Markdown or JSON while keeping tables and formulas. Comes with a CLI, a Python SDK and an agent skill.
An IBM library for document parsing: PDF, Office, HTML and images become one unified structure from which Markdown is easy to get. Supports vision models for hard pages.
A ready-made «chat with your documents» app: plug in any model and vector database, upload files and talk to them. Has a desktop version, Docker and agents.
A vector database for embedding search at large scale. It scales to cloud workloads, and for prototypes there is a lightweight local version, Milvus Lite.
Meta's classic library for fast similarity search and clustering of vectors. It is not a database but an index engine that many other systems are built on.
A lightweight alternative to GraphRAG: it builds a knowledge graph from documents and answers questions using both the graph and plain vector search. Ships with a server and a web UI.
A modular Microsoft system that builds an entity graph and community summaries from texts, so it can answer questions about a whole corpus rather than one passage.
A fast Rust vector database with payload filtering, hybrid search, quantization and a distributed mode. Available as a container and as a cloud service.
A long-term memory layer for agents: it remembers facts, links them into a graph and retrieves them on demand. The basic flow runs locally, even without an LLM key.
The easiest vector database to start with: it runs inside your app process, stores documents and finds similar ones. Great for prototypes, with a client-server mode and a cloud.
A Postgres extension that adds a vector type and nearest-neighbor search. It lets you keep embeddings next to regular data without running a separate database.
A vector database that stores both objects and vectors and combines semantic search with ordinary filters. It can call embedding models itself.
ETL for documents: it splits PDFs, emails, Word, HTML and images into typed elements (titles, paragraphs, tables) ready for chunking and embeddings.
An embedded search database, like SQLite for embeddings: it stores data in files and searches by keyword, vector and SQL. Handy for multimodal data.
A retrieval model that compares query and document token by token instead of with a single vector. That gives higher accuracy while searching large collections in tens of milliseconds.
Did we miss something?
Suggest a repository
Send a GitHub link and a few words on why it belongs here. Every suggestion is reviewed by hand — not everything gets in.










