github Actively maintained

peterk/warcworker

A dockerized, queued high fidelity web archiver based on Squidwarc

1 awesome list

Quick read

Stars
62
Forks
9
Open issues
6
Commits
34

Activity and growth

Latest capture 2026-09-03 03:04

Stars · last 7 days
0 0.0%
Commits · last 7 days
0 0.0%
Stars since tracking
0
Stored snapshots
7

Classification

Metadata

Language
Python
License
GPL-3.0
Default branch
master
Created
2018-07-21
First commit
2018-07-21
Last pushed
2024-07-09
GitHub updated
2026-02-08
Last synced
2026-09-03 03:04
Stack scanned
2026-09-03 03:04
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

7 observed captures since 2026-05-23. Observed captures are shown by default.

Stars from first capture 0

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

N0taN3rd/Squidwarc

Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head

178 stars
JavaScript 1 awesome list

N0taN3rd/node-warc

Parse And Create Web ARChive (WARC) files with node.js

105 stars
JavaScript 1 awesome list

netarchivesuite/solrwayback

A search interface and wayback machine for the UKWA Solr based warc-indexer framework.

147 stars
Java 1 awesome list

chfoo/warcat

Tool and library for handling Web ARChive (WARC) files.

165 stars
Python 1 awesome list

webrecorder/warcio

Streaming WARC/ARC library for fast web archive IO

471 stars
Python 1 awesome list