AIMLPM/markcrawl
Fast Python web crawler for RAG and AI ingestion. Extracts clean Markdown from any site for LLMs and vector stores.
馃殌馃 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
Appears on
Quick read
Latest capture 2026-09-04 04:09
4 paths
Agent instructions and tool configuration found in this repository.
Claude Code 4
135 observed captures since 2026-05-22. Observed captures are shown by default.
Stars from first capture +15144
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Fast Python web crawler for RAG and AI ingestion. Extracts clean Markdown from any site for LLMs and vector stores.
The context API to search, scrape, and interact with the web at scale. 馃敟
Transform developer documentation to clean Markdown
Crawlee鈥擜 web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Convert any URL to an LLM-friendly input with a simple prefix https://r.jina.ai/
The unified web layer for AI agents. Search (8 engines), stealth browse, auth, and act on 24 platforms. One npm install, self-hosted.