A small Microsoft utility that converts PDF, Word, PowerPoint, Excel and other files into Markdown that is easy to hand to a language model.
RAG & vector DBsunclecode
crawl4ai
An open-source Python crawler: it visits pages with a real browser and returns clean Markdown for LLMs. Runs locally, in Docker or as a cloud service.
Who it is for
For Python developers collecting website data for RAG and agents.
How to start
- Install:
pip install -U crawl4ai. - Install the browser once with
crawl4ai-setup, then verify withcrawl4ai-doctor. - Run
AsyncWebCrawler().arun(url=...)from the README example and printresult.markdown.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
More in «RAG & vector DBs»
A service that turns any website into clean Markdown or structured data for models. It can search, scrape, crawl a whole site and even click through a page.
A ready-made RAG engine with a UI: it parses documents with templates, chunks them, searches with citations and supports agentic retrieval. Deploys with Docker.
Parses complex PDFs, scans and Office files into Markdown or JSON while keeping tables and formulas. Comes with a CLI, a Python SDK and an agent skill.
An IBM library for document parsing: PDF, Office, HTML and images become one unified structure from which Markdown is easy to get. Supports vision models for hard pages.
A ready-made «chat with your documents» app: plug in any model and vector database, upload files and talk to them. Has a desktop version, Docker and agents.
