github Actively maintained

internetarchive/heritrix3

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.

1 awesome list

Quick read

Stars
3,313
Forks
792
Open issues
37
Commits
3,022

Activity and growth

Latest capture 2026-09-04 03:04

Stars · last 7 days
0 0.0%
Commits · last 7 days
0 0.0%
Stars since tracking
+86
Stored snapshots
7

Classification

Metadata

Language
Java
License
NOASSERTION
Default branch
master
Created
2011-10-21
First commit
2009-05-11
Last pushed
2026-09-02
GitHub updated
2026-09-03
Last synced
2026-09-04 03:04
Stack scanned
2026-09-04 03:04
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

7 observed captures since 2026-05-23. Observed captures are shown by default.

Stars from first capture +86

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

ukwa/ukwa-heritrix

The UKWA Heritrix3 custom modules and Docker builder.

11 stars
Java 1 awesome list

apache/incubator-heron

Apache Heron (Incubating) is a realtime, distributed, fault-tolerant stream processing engine from Twitter

3,627 stars
Java 1 awesome list

webrecorder/browsertrix

Browsertrix is the hosted, high-fidelity, browser-based crawling service from Webrecorder designed to make web archiving easier and more accessible for all!

468 stars
TypeScript 1 awesome list

webrecorder/browsertrix-crawler

Run a high-fidelity browser-based web archiving crawler in a single Docker container

1,127 stars
TypeScript 1 awesome list

herdrdev/herdr

the runtime your coding agents live on

28,899 stars
Rust 3 awesome lists