github Actively maintained

internetarchive/brozzler

brozzler - distributed browser-based web crawler

1 awesome list

Quick read

Stars
812
Forks
114
Open issues
63
Commits
1,792

Activity and growth

Latest capture 2026-07-31 03:04

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+16
Stored snapshots
6

Classification

Technology stack

Metadata

Language
Python
License
Apache-2.0
Default branch
master
Created
2015-07-13
First commit
2014-01-21
Last pushed
2026-07-27
GitHub updated
2026-07-30
Last synced
2026-07-31 03:04
Stack scanned
2026-07-31 03:04
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-23. Observed captures are shown by default.

Stars from first capture +16

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

s0rg/crawley

The unix-way web crawler

341 stars
Go 2 awesome lists

webrecorder/browsertrix

Browsertrix is the hosted, high-fidelity, browser-based crawling service from Webrecorder designed to make web archiving easier and more accessible for all!

468 stars
TypeScript 1 awesome list

webrecorder/browsertrix-crawler

Run a high-fidelity browser-based web archiving crawler in a single Docker container

1,127 stars
TypeScript 1 awesome list

ArchiveTeam/grab-site

The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns

1,602 stars
Python 1 awesome list

openzim/zimit

Make a ZIM file from any Web site and surf offline!

827 stars
Python 1 awesome list