A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
RAG & vector DBsfacebookresearch
faiss
Meta's classic library for fast similarity search and clustering of vectors. It is not a database but an index engine that many other systems are built on.
Who it is for
For those who need fast vector search inside their own code without a separate server.
How to start
- Install the prebuilt
faiss-cpu(orfaiss-gpu) via Anaconda, as the Installing section describes. - To build from source, open
INSTALL.md: you need cmake and a BLAS implementation. - Follow the full documentation and repeat the index-building example.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
A library for efficient similarity search and clustering of dense vectors.
More in «RAG & vector DBs»
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A ready-made RAG engine with a UI: it parses documents with templates, chunks them, searches with citations and supports agentic retrieval. Deploys with Docker.
An open-source Python crawler: it visits pages with a real browser and returns clean Markdown for LLMs. Runs locally, in Docker or as a cloud service.
Parses complex PDFs, scans and Office files into Markdown or JSON while keeping tables and formulas. Comes with a CLI, a Python SDK and an agent skill.
An IBM library for document parsing: PDF, Office, HTML and images become one unified structure from which Markdown is easy to get. Supports vision models for hard pages.
