Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
A curated list of awesome big data frameworks, ressources and other awesomeness.
GitHub stars and default-branch commits for oxnr/awesome-bigdata.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Apache Superset is a Data Visualization and Data Exploration Platform
Deep Learning for humans
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
Apache Spark - A unified analytics engine for large-scale data processing
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Make Your Company Data Driven. Connect to any data source, easily visualize, dashboard and share your data.
High-performance, scalable time-series database designed for Industrial IoT (IIoT) scenarios
Data Apps & Dashboards for Python. No JavaScript Required.
matplotlib: plotting with Python
Luigi is a Python module that helps you build complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization etc. It also comes with Hadoop support built in.
An orchestration platform for the development, production, and observation of data assets.
A lightweight, lightning-fast, in-process vector database
YugabyteDB - the cloud native distributed SQL database for mission-critical applications.
Easy & Flexible Alerting With ElasticSearch
The Open Source Feature Store for AI/ML
The live data layer for apps and AI agents. Create up-to-the-second views into your business, just using SQL
Numenta Platform for Intelligent Computing is an implementation of Hierarchical Temporal Memory (HTM), a theory of intelligence based strictly on the neuroscience of the neocortex.
Aim 💫 — An easy-to-use & supercharged open-source experiment tracker.
The AI-native database built for LLM applications, providing incredibly fast hybrid search of dense vector, sparse vector, tensor (multi-vector), and full-text.
Deploy and manage containers (including Docker) on top of Apache Mesos at scale.
chDB is an in-process OLAP SQL Engine 🚀 powered by ClickHouse
A lightweight opinionated ETL framework, halfway between plain scripts and Apache Airflow
Serverless proxy for Spark cluster
A distributed Spark/Scala implementation of the isolation forest and extended isolation forest algorithms for unsupervised outlier detection, featuring support for scalable training and ONNX export for easy cross-platform inference.