github Actively maintained

ArchiveTeam/grab-site

The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns

1 awesome list

Quick read

Stars
1,602
Forks
157
Open issues
103
Commits
1,174

Activity and growth

Latest capture 2026-07-31 03:03

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+31
Stored snapshots
6

Classification

Metadata

Language
Python
License
NOASSERTION
Default branch
master
Created
2015-02-05
First commit
2015-02-05
Last pushed
2025-05-23
GitHub updated
2026-07-29
Last synced
2026-07-31 03:03
Stack scanned
2026-07-31 03:03
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-23. Observed captures are shown by default.

Stars from first capture +31

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

N0taN3rd/Squidwarc

Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head

178 stars
JavaScript 1 awesome list

ArchiveTeam/wpull

Wget-compatible web downloader and crawler.

613 stars
HTML 1 awesome list

s0rg/crawley

The unix-way web crawler

341 stars
Go 2 awesome lists

netarchivesuite/solrwayback

A search interface and wayback machine for the UKWA Solr based warc-indexer framework.

147 stars
Java 1 awesome list

harvard-lil/scoop

🍨 High-fidelity, browser-based, single-page web archiving library and CLI for witnessing the web.

210 stars
JavaScript 1 awesome list