The AI movement's repository database

The best AI repositories

224 living projects — from agent frameworks to prompt-injection defense. Each comes with our own note: why it matters, who it is for and how to start in three steps.

Stars, forks and activity refresh daily via the GitHub API · updated October 6, 2026

repositories
224
categories
11
stars in total
11M
new stars in 30 days
+235K

Last 30 days

Growing fastest

The most new stars over the last 30 days.

Full growth ranking
Growing fastestRookie of the month

A skills set and methodology for coding agents: clarify the task, plan, write tests, then code. The agent works with more discipline.

296Kstars+15K in 30 dShell

firecrawl

firecrawl

Growing fastest

A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.

189Kstars+12K in 30 dTypeScript

microsoft

markitdown

Growing fastest

A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.

189Kstars+11K in 30 dPython

anomalyco

opencode

Growing fastest

An open coding agent for the terminal with a desktop app. It is not tied to one provider: plug in whichever models you have.

212Kstars+7.6K in 30 dTypeScript

github

spec-kit

Growing fastest

A toolkit for spec-driven development: principles and a spec first, then a plan and tasks, and only then code. Works with several coding agents.

140Kstars+7.1K in 30 dPython

openai

codex

Growing fastest

OpenAI's lightweight coding agent that runs in the terminal and works on your project's code. Written in Rust.

128Kstars+6.8K in 30 dRust

Categories

Agent frameworksFrameworks and orchestration: what agents and multi-agent systems are built from.23 repositoriesCoding agents & CLIAgents that write code in the terminal and editor, and the methods for working with them.21 repositoriesMCPThe Model Context Protocol spec, official SDKs and the best servers.20 repositoriesOpen models & inferenceOpen weights, local runs, fast inference and fine-tuning.38 repositoriesRAG & vector DBsSearch over your own data: document parsing, crawlers, vector databases, knowledge graphs.19 repositoriesEvals & observabilityHow to measure a model or agent and see what happens inside it.17 repositoriesPrompting & contextPrompts, structured output, skills and context engineering.16 repositoriesVoice, video, multimodalSpeech recognition and synthesis, voice agents, images and video.19 repositoriesAutomationNo-code workflows, browser and computer-use agents.15 repositoriesAwesome lists & learningCourses, books and awesome lists worth starting with.21 repositoriesSecurityPrompt injection, guardrails, red teaming and agent scanners.15 repositoriesMissing your favorite?Suggest a repository — we will review it and add it if it earns a spot.Suggest a repo →

Catalog

19 repositories

microsoft

markitdown

Growing fastest

A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.

189Kstars+11K in 30 d14KPythonCommit 2 days ago

firecrawl

firecrawl

Growing fastest

A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.

189Kstars+12K in 30 d10KTypeScriptCommit 1 day ago

infiniflow

ragflow

A ready-made RAG engine with a UI: it parses documents with templates, chunks them, searches with citations and supports agentic retrieval. Deploys with Docker.

92Kstars+1.8K in 30 d11KGoCommit 1 day ago

unclecode

crawl4ai

An open-source Python crawler: it visits pages with a real browser and returns clean Markdown for LLMs. Runs locally, in Docker or as a cloud service.

85Kstars+3.2K in 30 d8.8KPythonCommit 1 day ago

opendatalab

MinerU

Parses complex PDFs, scans and Office files into Markdown or JSON while keeping tables and formulas. Comes with a CLI, a Python SDK and an agent skill.

81Kstars+2K in 30 d6.8KPythonCommit today

docling-project

docling

An IBM library for document parsing: PDF, Office, HTML and images become one unified structure from which Markdown is easy to get. Supports vision models for hard pages.

68Kstars+2.5K in 30 d5KPythonCommit 1 day ago

Mintplex-Labs

anything-llm

A ready-made «chat with your documents» app: plug in any model and vector database, upload files and talk to them. Has a desktop version, Docker and agents.

67Kstars+1.4K in 30 d7.4KJavaScriptCommit 2 days ago

milvus-io

milvus

A vector database for embedding search at large scale. It scales to cloud workloads, and for prototypes there is a lightweight local version, Milvus Lite.

46Kstars+365 in 30 d4.3KGoCommit today

facebookresearch

faiss

Meta's classic library for fast similarity search and clustering of vectors. It is not a database but an index engine that many other systems are built on.

41Kstars+238 in 30 d4.6KC++Commit 3 days ago

A lightweight alternative to GraphRAG: it builds a knowledge graph from documents and answers questions using both the graph and plain vector search. Ships with a server and a web UI.

40Kstars+615 in 30 d5.7KPythonCommit 3 days ago

microsoft

graphrag

A modular Microsoft system that builds an entity graph and community summaries from texts, so it can answer questions about a whole corpus rather than one passage.

36Kstars+424 in 30 d3.8KPythonCommit 1 day ago

qdrant

qdrant

A fast Rust vector database with payload filtering, hybrid search, quantization and a distributed mode. Available as a container and as a cloud service.

35Kstars+588 in 30 d2.7KRustCommit today

topoteretes

cognee

A long-term memory layer for agents: it remembers facts, links them into a graph and retrieves them on demand. The basic flow runs locally, even without an LLM key.

31Kstars+984 in 30 d3.2KPythonCommit 1 day ago

chroma-core

chroma

The easiest vector database to start with: it runs inside your app process, stores documents and finds similar ones. Great for prototypes, with a client-server mode and a cloud.

29Kstars+243 in 30 d2.5KRustCommit 2 days ago

pgvector

pgvector

A Postgres extension that adds a vector type and nearest-neighbor search. It lets you keep embeddings next to regular data without running a separate database.

23Kstars+347 in 30 d1.3KCCommit 4 days ago

weaviate

weaviate

A vector database that stores both objects and vectors and combines semantic search with ordinary filters. It can call embedding models itself.

17Kstars+91 in 30 d1.4KGoCommit today

Unstructured-IO

unstructured

ETL for documents: it splits PDFs, emails, Word, HTML and images into typed elements (titles, paragraphs, tables) ready for chunking and embeddings.

16Kstars+147 in 30 d1.4KHTMLCommit today

lancedb

lancedb

An embedded search database, like SQLite for embeddings: it stores data in files and searches by keyword, vector and SQL. Handy for multimodal data.

12Kstars+258 in 30 d1.1KRustCommit 1 day ago

stanford-futuredata

ColBERT

A retrieval model that compares query and document token by token instead of with a single vector. That gives higher accuracy while searching large collections in tens of milliseconds.

3.9Kstars+20 in 30 d476PythonCommit 12 months ago

Did we miss something?

Suggest a repository

Send a GitHub link and a few words on why it belongs here. Every suggestion is reviewed by hand — not everything gets in.