github Actively maintained

apache/hudi

Upserts, Deletes And Incremental Processing on Big Data.

1 awesome list

Quick read

Stars
6,203
Forks
2,495
Open issues
2,905
Commits
7,705

Activity and growth

Latest capture 2026-08-02 03:08

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+38
Stored snapshots
6

Classification

Metadata

Language
Java
License
Apache-2.0
Default branch
master
Created
2016-12-14
First commit
2016-12-16
Last pushed
2026-08-01
GitHub updated
2026-08-01
Last synced
2026-08-02 03:08
Stack scanned
2026-08-02 03:08
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +38

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

apache/spark

Apache Spark - A unified analytics engine for large-scale data processing

43,763 stars
Scala 4 awesome lists

apache/flink

Apache Flink

26,233 stars
Java 1 awesome list

TIBCOSoftware/snappydata

Project SnappyData - memory optimized analytics database, based on Apache Spark™ and Apache Geode™. Stream, Transact, Analyze, Predict in one cluster

1,033 stars
Scala 1 awesome list

h2oai/h2o-3

H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.

7,499 stars
Jupyter Notebook 3 awesome lists

apache/druid

Apache Druid: a high performance real-time analytics database.

14,050 stars
Java 3 awesome lists

apache/gobblin

A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

2,271 stars
Java 1 awesome list