A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
RAG & vector DBsmicrosoft
graphrag
A modular Microsoft system that builds an entity graph and community summaries from texts, so it can answer questions about a whole corpus rather than one passage.
Who it is for
For analysts and developers who need answers over a large body of text as a whole.
How to start
- Open the command-line quickstart in the docs at microsoft.github.io/graphrag.
- Read the Prompt Tuning section to adapt prompts to your data.
- Read the Responsible AI FAQ from the README before running on real data.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
A modular graph-based Retrieval-Augmented Generation (RAG) system
More in «RAG & vector DBs»
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A ready-made RAG engine with a UI: it parses documents with templates, chunks them, searches with citations and supports agentic retrieval. Deploys with Docker.
An open-source Python crawler: it visits pages with a real browser and returns clean Markdown for LLMs. Runs locally, in Docker or as a cloud service.
Parses complex PDFs, scans and Office files into Markdown or JSON while keeping tables and formulas. Comes with a CLI, a Python SDK and an agent skill.
An IBM library for document parsing: PDF, Office, HTML and images become one unified structure from which Markdown is easy to get. Supports vision models for hard pages.
