The AI movement's repository database

The best AI repositories

224 living projects — from agent frameworks to prompt-injection defense. Each comes with our own note: why it matters, who it is for and how to start in three steps.

Stars, forks and activity refresh daily via the GitHub API · updated October 6, 2026

repositories
224
categories
11
stars in total
11M
new stars in 30 days
+235K

Last 30 days

Growing fastest

The most new stars over the last 30 days.

Full growth ranking
Growing fastestRookie of the month

A skills set and methodology for coding agents: clarify the task, plan, write tests, then code. The agent works with more discipline.

296Kstars+15K in 30 dShell

firecrawl

firecrawl

Growing fastest

A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.

189Kstars+12K in 30 dTypeScript

microsoft

markitdown

Growing fastest

A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.

189Kstars+11K in 30 dPython

anomalyco

opencode

Growing fastest

An open coding agent for the terminal with a desktop app. It is not tied to one provider: plug in whichever models you have.

212Kstars+7.6K in 30 dTypeScript

github

spec-kit

Growing fastest

A toolkit for spec-driven development: principles and a spec first, then a plan and tasks, and only then code. Works with several coding agents.

140Kstars+7.1K in 30 dPython

openai

codex

Growing fastest

OpenAI's lightweight coding agent that runs in the terminal and works on your project's code. Written in Rust.

128Kstars+6.8K in 30 dRust

Categories

Agent frameworksFrameworks and orchestration: what agents and multi-agent systems are built from.23 repositoriesCoding agents & CLIAgents that write code in the terminal and editor, and the methods for working with them.21 repositoriesMCPThe Model Context Protocol spec, official SDKs and the best servers.20 repositoriesOpen models & inferenceOpen weights, local runs, fast inference and fine-tuning.38 repositoriesRAG & vector DBsSearch over your own data: document parsing, crawlers, vector databases, knowledge graphs.19 repositoriesEvals & observabilityHow to measure a model or agent and see what happens inside it.17 repositoriesPrompting & contextPrompts, structured output, skills and context engineering.16 repositoriesVoice, video, multimodalSpeech recognition and synthesis, voice agents, images and video.19 repositoriesAutomationNo-code workflows, browser and computer-use agents.15 repositoriesAwesome lists & learningCourses, books and awesome lists worth starting with.21 repositoriesSecurityPrompt injection, guardrails, red teaming and agent scanners.15 repositoriesMissing your favorite?Suggest a repository — we will review it and add it if it earns a spot.Suggest a repo →

Catalog

14 repositories

langfuse

langfuse

An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.

35Kstars+1.2K in 30 d3.9KTypeScriptCommit 1 day ago

mastra-ai

mastra

A TypeScript framework for agents and AI apps: agents, workflows, memory and quality evaluation in one place. A good fit for Node-based web teams.

29Kstars+904 in 30 d2.9KTypeScriptCommit today

trycua

cua

Growing fastest

Tooling for computer-use agents: drivers for several OSes, virtual machines and benchmarks for training and evaluation.

28Kstars+6.2K in 30 d2KRustCommit today

mlflow

mlflow

A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.

28Kstars+455 in 30 d6.4KPythonCommit 1 day ago

promptfoo

promptfoo

A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.

26Kstars+919 in 30 d2.4KTypeScriptCommit 1 day ago

comet-ml

opik

A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.

22Kstars+618 in 30 d1.8KPythonCommit today

Google's code-first agent framework: build agents in Python, run evaluations and deploy. Tuned for Gemini but works with other models too.

22Kstars+326 in 30 d4.1KPythonCommit 1 day ago

confident-ai

deepeval

A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.

19Kstars+533 in 30 d2KPythonCommit 3 days ago

vibrantlabsai

ragas

A metrics library for evaluating RAG and other LLM apps: it checks how well an answer relies on retrieved sources. Can also generate test sets.

16Kstars+322 in 30 d1.7KPythonCommit 7 months ago

The standard toolkit for running models through hundreds of academic benchmarks. Many open model leaderboards are built on it.

14Kstars+242 in 30 d3.6KPythonCommit 22 days ago

Arize-ai

phoenix

An observability and evaluation tool for LLMs: starts locally with one command and shows traces, experiments and datasets. Built on OpenTelemetry.

12Kstars+403 in 30 d1.2KPythonCommit today

SWE-bench

SWE-bench

The main benchmark for coding agents: a model must fix real issues from GitHub projects, and the patch is verified with tests in containers.

6Kstars+197 in 30 d1KPythonCommit 18 days ago

UKGovernmentBEIS

inspect_ai

A model-evaluation framework from the UK AI Security Institute: tool use, multi-turn dialog, model-graded scoring and over 200 ready-made evals.

2.9Kstars+227 in 30 d773PythonCommit 1 day ago

harbor-framework

terminal-bench-1

A benchmark of hard terminal tasks: a set of assignments and a runner where the agent acts as an admin or developer. Has a leaderboard.

2.6Kstars572PythonCommit 3 months ago

Did we miss something?

Suggest a repository

Send a GitHub link and a few words on why it belongs here. Every suggestion is reviewed by hand — not everything gets in.