The AI movement's repository database

The best AI repositories

224 living projects — from agent frameworks to prompt-injection defense. Each comes with our own note: why it matters, who it is for and how to start in three steps.

Stars, forks and activity refresh daily via the GitHub API · updated October 6, 2026

repositories
224
categories
11
stars in total
11M
new stars in 30 days
+235K

Last 30 days

Growing fastest

The most new stars over the last 30 days.

Full growth ranking
Growing fastestRookie of the month

A skills set and methodology for coding agents: clarify the task, plan, write tests, then code. The agent works with more discipline.

296Kstars+15K in 30 dShell

firecrawl

firecrawl

Growing fastest

A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.

189Kstars+12K in 30 dTypeScript

microsoft

markitdown

Growing fastest

A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.

189Kstars+11K in 30 dPython

anomalyco

opencode

Growing fastest

An open coding agent for the terminal with a desktop app. It is not tied to one provider: plug in whichever models you have.

212Kstars+7.6K in 30 dTypeScript

github

spec-kit

Growing fastest

A toolkit for spec-driven development: principles and a spec first, then a plan and tasks, and only then code. Works with several coding agents.

140Kstars+7.1K in 30 dPython

openai

codex

Growing fastest

OpenAI's lightweight coding agent that runs in the terminal and works on your project's code. Written in Rust.

128Kstars+6.8K in 30 dRust

Categories

Agent frameworksFrameworks and orchestration: what agents and multi-agent systems are built from.23 repositoriesCoding agents & CLIAgents that write code in the terminal and editor, and the methods for working with them.21 repositoriesMCPThe Model Context Protocol spec, official SDKs and the best servers.20 repositoriesOpen models & inferenceOpen weights, local runs, fast inference and fine-tuning.38 repositoriesRAG & vector DBsSearch over your own data: document parsing, crawlers, vector databases, knowledge graphs.19 repositoriesEvals & observabilityHow to measure a model or agent and see what happens inside it.17 repositoriesPrompting & contextPrompts, structured output, skills and context engineering.16 repositoriesVoice, video, multimodalSpeech recognition and synthesis, voice agents, images and video.19 repositoriesAutomationNo-code workflows, browser and computer-use agents.15 repositoriesAwesome lists & learningCourses, books and awesome lists worth starting with.21 repositoriesSecurityPrompt injection, guardrails, red teaming and agent scanners.15 repositoriesMissing your favorite?Suggest a repository — we will review it and add it if it earns a spot.Suggest a repo →

Catalog

17 repositories

langfuse

langfuse

An open platform that shows every model call: traces, cost, quality scores and prompt versions. Use it in the cloud or self-host it.

35Kstars+1.2K in 30 d3.9KTypeScriptCommit 1 day ago

mlflow

mlflow

A long-standing ML platform that now covers LLM apps too: tracing, evaluation, a prompt registry and a model gateway. Plugs into most frameworks.

28Kstars+455 in 30 d6.4KPythonCommit 1 day ago

promptfoo

promptfoo

A command-line tool for testing prompts and agents: describe cases in a config and compare models. Also does red teaming and vulnerability scanning. Plugs into CI.

26Kstars+919 in 30 d2.4KTypeScriptCommit 1 day ago

comet-ml

opik

A platform for debugging and evaluating LLM apps: traces, automated quality checks and dashboards. Supports RAG and agent chains, and can be self-hosted.

22Kstars+618 in 30 d1.8KPythonCommit today

openai

evals

OpenAI's framework for evaluating models plus a registry of ready-made benchmarks. A historically important project that shows how eval templates are structured.

20Kstars+181 in 30 d3.1KPythonCommit 6 months ago

confident-ai

deepeval

A pytest-style framework for testing LLM apps: ready-made metrics for RAG, agents and chatbots, plus trace-based checks. Results can be pushed to a cloud.

19Kstars+533 in 30 d2KPythonCommit 3 days ago

vibrantlabsai

ragas

A metrics library for evaluating RAG and other LLM apps: it checks how well an answer relies on retrieved sources. Can also generate test sets.

16Kstars+322 in 30 d1.7KPythonCommit 7 months ago

The standard toolkit for running models through hundreds of academic benchmarks. Many open model leaderboards are built on it.

14Kstars+242 in 30 d3.6KPythonCommit 22 days ago

Arize-ai

phoenix

An observability and evaluation tool for LLMs: starts locally with one command and shows traces, experiments and datasets. Built on OpenTelemetry.

12Kstars+403 in 30 d1.2KPythonCommit today

traceloop

openllmetry

An OpenTelemetry-based toolkit that automatically traces calls to models, vector databases and frameworks. Data can go to any compatible monitoring system.

7.5Kstars+74 in 30 d1.1KPythonCommit 2 days ago

Helicone

helicone

A gateway and observability platform: change the API address in your code and all model requests get logged, costed and compared. Can be self-hosted.

6.2Kstars+71 in 30 d682TypeScriptCommit 19 days ago

SWE-bench

SWE-bench

The main benchmark for coding agents: a model must fix real issues from GitHub projects, and the patch is verified with tests in containers.

6Kstars+197 in 30 d1KPythonCommit 18 days ago

AgentOps-AI

agentops

An SDK for observing agents: it records sessions, tracks cost and shows the chain of steps. Integrates with CrewAI, LangChain, OpenAI Agents SDK and others.

5.9Kstars642PythonCommit 3 months ago

fchollet

ARC-AGI

A famous abstract-reasoning test: 800 puzzles on coloured grids where you infer the rule from a few examples. Easy for people, long out of reach for models. The repo holds the data and a page for solving tasks by hand.

4.8Kstars+19 in 30 d726JavaScriptCommit 1 year ago

pydantic

logfire

An observability platform from the makers of Pydantic: it traces plain code, model calls and agents. Built on OpenTelemetry, with an open-source SDK.

4.5Kstars+53 in 30 d304PythonCommit 1 day ago

UKGovernmentBEIS

inspect_ai

A model-evaluation framework from the UK AI Security Institute: tool use, multi-turn dialog, model-graded scoring and over 200 ready-made evals.

2.9Kstars+227 in 30 d773PythonCommit 1 day ago

harbor-framework

terminal-bench-1

A benchmark of hard terminal tasks: a set of assignments and a runner where the agent acts as an admin or developer. Has a leaderboard.

2.6Kstars572PythonCommit 3 months ago

Did we miss something?

Suggest a repository

Send a GitHub link and a few words on why it belongs here. Every suggestion is reviewed by hand — not everything gets in.