github Actively maintained

apache/spark

Apache Spark - A unified analytics engine for large-scale data processing

Quick read

Stars
43,763
Forks
29,298
Open issues
465
Commits
49,379

Activity and growth

Latest capture 2026-08-02 03:10

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+443
Stored snapshots
11

Classification

Technology stack

Metadata

Language
Scala
License
Apache-2.0
Default branch
master
Created
2014-02-25
First commit
2010-03-29
Last pushed
2026-08-01
GitHub updated
2026-08-02
Last synced
2026-08-02 03:10
Stack scanned
2026-08-02 03:10
Archived
No

AI development signals

2 paths

Agent instructions and tool configuration found in this repository.

Agent instructions

Claude Code

View paths

Growth history

Tracked growth

11 observed captures since 2026-05-22. Observed captures are shown by default.

Stars from first capture +443

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

apache/hudi

Upserts, Deletes And Incremental Processing on Big Data.

6,203 stars
Java 1 awesome list

apache/maven

Apache Maven core

5,338 stars
Java 1 awesome list

TIBCOSoftware/snappydata

Project SnappyData - memory optimized analytics database, based on Apache Spark™ and Apache Geode™. Stream, Transact, Analyze, Predict in one cluster

1,033 stars
Scala 1 awesome list