N0taN3rd/Squidwarc
Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
Appears on
Quick read
Latest capture 2026-07-31 03:03
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture +31
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head
Wget-compatible web downloader and crawler.
The unix-way web crawler
A search interface and wayback machine for the UKWA Solr based warc-indexer framework.
A whirlwind tour of Common Crawl's data using Python
🍨 High-fidelity, browser-based, single-page web archiving library and CLI for witnessing the web.