All repositories

RAG & vector DBsUnstructured-IO

unstructured

ETL for documents: it splits PDFs, emails, Word, HTML and images into typed elements (titles, paragraphs, tables) ready for chunking and embeddings.

Who it is for

For engineers building a document-preparation pipeline for LLMs.

How to start

  1. Install: pip install "unstructured[all-docs]" (plain text only needs pip install unstructured).
  2. Install system dependencies: libmagic-dev, poppler-utils, tesseract-ocr.
  3. Or run the prebuilt container: docker pull downloads.unstructured.io/unstructured-io/unstructured:latest.

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+147Sep 6 — Oct 6
15,38215,529

Author's description

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

microsoft

markitdown

Growing fastest

A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.

189Kstars+11K in 30 dPython

firecrawl

firecrawl

Growing fastest

A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.

189Kstars+12K in 30 dTypeScript

infiniflow

ragflow

A ready-made RAG engine with a UI: it parses documents with templates, chunks them, searches with citations and supports agentic retrieval. Deploys with Docker.

92Kstars+1.8K in 30 dGo

unclecode

crawl4ai

An open-source Python crawler: it visits pages with a real browser and returns clean Markdown for LLMs. Runs locally, in Docker or as a cloud service.

85Kstars+3.2K in 30 dPython

opendatalab

MinerU

Parses complex PDFs, scans and Office files into Markdown or JSON while keeping tables and formulas. Comes with a CLI, a Python SDK and an agent skill.

81Kstars+2K in 30 dPython

docling-project

docling

An IBM library for document parsing: PDF, Office, HTML and images become one unified structure from which Markdown is easy to get. Supports vision models for hard pages.

68Kstars+2.5K in 30 dPython