github Actively maintained

internetarchive/Sparkling

Internet Archive's Sparkling Data Processing Library

1 awesome list

Quick read

Stars
17
Forks
2
Open issues
1
Commits
47

Activity and growth

Latest capture 2026-09-03 03:04

Stars · last 7 days
0 0.0%
Commits · last 7 days
0 0.0%
Stars since tracking
+1
Stored snapshots
7

Metadata

Language
Scala
License
MIT
Default branch
main
Created
2022-04-28
First commit
2022-04-28
Last pushed
2026-05-04
GitHub updated
2026-06-08
Last synced
2026-09-03 03:04
Stack scanned
2026-09-03 03:04
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

7 observed captures since 2026-05-23. Observed captures are shown by default.

Stars from first capture +1

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

helgeho/ArchiveSpark

An Apache Spark framework for easy data processing, extraction as well as derivation for web archives and archival collections, developed at Internet Archive.

162 stars
Scala 1 awesome list

archivesunleashed/aut

The Archives Unleashed Toolkit is an open-source toolkit for analyzing web archives.

158 stars
Scala 1 awesome list

svenkreiss/pysparkling

A pure Python implementation of Apache Spark's RDD and DStream interfaces.

270 stars
Python 1 awesome list

TIBCOSoftware/snappydata

Project SnappyData - memory optimized analytics database, based on Apache Spark™ and Apache Geode™. Stream, Transact, Analyze, Predict in one cluster

1,033 stars
Scala 1 awesome list

internetarchive/arch

Web application for distributed compute analysis of Archive-It web archive collections.

20 stars
Scala 1 awesome list