github Actively maintained

helgeho/Web2Warc

An easy-to-use and highly customizable crawler that enables you to create your own little Web archives (WARC/CDX)

1 awesome list

Quick read

Stars
26
Forks
3
Open issues
1
Commits
11

Activity and growth

Latest capture 2026-09-04 03:04

Stars · last 7 days
0 0.0%
Commits · last 7 days
0 0.0%
Stars since tracking
0
Stored snapshots
7

Metadata

Language
Scala
License
MIT
Default branch
master
Created
2016-01-29
First commit
2016-01-29
Last pushed
2017-10-09
GitHub updated
2026-05-21
Last synced
2026-09-04 03:04
Stack scanned
2026-09-04 03:04
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

7 observed captures since 2026-05-23. Observed captures are shown by default.

Stars from first capture 0

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

nla/httrack2warc

Converts HTTrack crawls to WARC files

34 stars
Java 1 awesome list

ukwa/webarchive-discovery

Please note that the warc-indexer tool & code is now supported by NetArchiveSuite. The 'warc-indexer' directory and code that exists in this repo is now only for reference. For support and issues of 'warc-indexer', please communicate with NetArchiveSuite.

133 stars
Java 1 awesome list

N0taN3rd/Squidwarc

Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head

178 stars
JavaScript 1 awesome list

webrecorder/warcio

Streaming WARC/ARC library for fast web archive IO

471 stars
Python 1 awesome list