s0rg/crawley
The unix-way web crawler
Extract web archive data using Wayback Machine and Common Crawl
Appears on
Quick read
Latest capture 2026-09-03 03:04
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
7 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture +4
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
The unix-way web crawler
Easy-to-use Web archiver
DuckDB extension to fetch pages from Wayback Machine & Common Crawl
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
A whirlwind tour of Common Crawl's data using Python