ArchiveTeam/grab-site
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
Wget-compatible web downloader and crawler.
Appears on
Quick read
Latest capture 2026-07-31 03:04
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture +5
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
Tools for downloading and preserving wikis. We archive wikis, from Wikipedia to tiniest wikis. As of 2026, WikiTeam has preserved more than 600,000 wikis.
Wget with Lua extension
An archiving tool with an IM-style interface that prioritizes privacy and accessibility, integrated with various archival services including Internet Archive, archive.today, Ghostarchive, IPFS, Telegraph, and file systems.
Core Python Web Archiving Toolkit for replay and recording of web archives
Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages.