github Actively maintained

N0taN3rd/Squidwarc

Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head

1 awesome list

Quick read

Stars
178
Forks
25
Open issues
11
Commits
119

Activity and growth

Latest capture 2026-07-31 03:05

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+4
Stored snapshots
6

Classification

Metadata

Language
JavaScript
License
Apache-2.0
Default branch
master
Created
2017-07-20
First commit
2017-07-20
Last pushed
2020-05-19
GitHub updated
2026-07-08
Last synced
2026-07-31 03:05
Stack scanned
2026-07-31 03:05
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-23. Observed captures are shown by default.

Stars from first capture +4

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

peterk/warcworker

A dockerized, queued high fidelity web archiver based on Squidwarc

62 stars
Python 1 awesome list

N0taN3rd/node-warc

Parse And Create Web ARChive (WARC) files with node.js

105 stars
JavaScript 1 awesome list

ArchiveTeam/grab-site

The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns

1,602 stars
Python 1 awesome list

helgeho/Web2Warc

An easy-to-use and highly customizable crawler that enables you to create your own little Web archives (WARC/CDX)

26 stars
Scala 1 awesome list

s0rg/crawley

The unix-way web crawler

341 stars
Go 2 awesome lists

netarchivesuite/solrwayback

A search interface and wayback machine for the UKWA Solr based warc-indexer framework.

147 stars
Java 1 awesome list