All repositories

RAG & vector DBsHKUDS

LightRAG

A lightweight alternative to GraphRAG: it builds a knowledge graph from documents and answers questions using both the graph and plain vector search. Ships with a server and a web UI.

Who it is for

For those who want graph RAG that is simpler and cheaper than heavyweight options.

How to start

  1. Install the server: uv tool install "lightrag-hku[api]".
  2. Download env.example from the repo and turn it into .env with your keys.
  3. Launch lightrag-server, but configure authentication in .env before exposing it to a network.

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+615Sep 6 — Oct 5
39,36739,982

Author's description

[EMNLP2025] LightRAG: Simple and Fast Retrieval-Augmented Generation

microsoft

markitdown

Growing fastest

A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.

189Kstars+11K in 30 dPython

firecrawl

firecrawl

Growing fastest

A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.

189Kstars+12K in 30 dTypeScript

infiniflow

ragflow

A ready-made RAG engine with a UI: it parses documents with templates, chunks them, searches with citations and supports agentic retrieval. Deploys with Docker.

92Kstars+1.8K in 30 dGo

unclecode

crawl4ai

An open-source Python crawler: it visits pages with a real browser and returns clean Markdown for LLMs. Runs locally, in Docker or as a cloud service.

85Kstars+3.2K in 30 dPython

opendatalab

MinerU

Parses complex PDFs, scans and Office files into Markdown or JSON while keeping tables and formulas. Comes with a CLI, a Python SDK and an agent skill.

81Kstars+2K in 30 dPython

docling-project

docling

An IBM library for document parsing: PDF, Office, HTML and images become one unified structure from which Markdown is easy to get. Supports vision models for hard pages.

68Kstars+2.5K in 30 dPython