N0taN3rd/Squidwarc
Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head
A dockerized, queued high fidelity web archiver based on Squidwarc
Appears on
Quick read
Latest capture 2026-09-03 03:04
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
7 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head
Parse And Create Web ARChive (WARC) files with node.js
A search interface and wayback machine for the UKWA Solr based warc-indexer framework.
Tool and library for handling Web ARChive (WARC) files.
Streaming WARC/ARC library for fast web archive IO
WarcDB: Web crawl data as SQLite databases.