github Actively maintained

helgeho/ArchiveSpark

An Apache Spark framework for easy data processing, extraction as well as derivation for web archives and archival collections, developed at Internet Archive.

1 awesome list

Quick read

Stars
162
Forks
19
Open issues
5
Commits
154

Activity and growth

Latest capture 2026-09-03 03:03

Stars · last 7 days
0 0.0%
Commits · last 7 days
0 0.0%
Stars since tracking
+2
Stored snapshots
7

Classification

Metadata

Language
Scala
License
MIT
Default branch
master
Created
2015-08-06
First commit
2015-08-06
Last pushed
2025-10-08
GitHub updated
2026-08-05
Last synced
2026-09-03 03:03
Stack scanned
2026-09-03 03:03
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

7 observed captures since 2026-05-23. Observed captures are shown by default.

Stars from first capture +2

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

archivesunleashed/aut

The Archives Unleashed Toolkit is an open-source toolkit for analyzing web archives.

158 stars
Scala 1 awesome list

helgeho/Web2Warc

An easy-to-use and highly customizable crawler that enables you to create your own little Web archives (WARC/CDX)

26 stars
Scala 1 awesome list

archivesunleashed/twut

An open-source toolkit for analyzing line-oriented JSON Twitter archives with Apache Spark.

10 stars
Scala 1 awesome list

internetarchive/arch

Web application for distributed compute analysis of Archive-It web archive collections.

20 stars
Scala 1 awesome list