Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
OpenRefine is a free, open source power tool for working with messy data and improving it
Rich MySQL Terminal Client with AutoCompletion, Syntax Highlighting, and Dataframes
Always know what to expect from your data.
Low-code framework for building custom LLMs, neural networks, and other AI models
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
Statsmodels: statistical modeling and econometrics in Python
Always know what to expect from your data.
Fast and Accurate ML in 3 Lines of Code
Modin: Scale your Pandas workflows by changing a single line of code
A Python Automated Machine Learning tool that optimizes machine learning pipelines using genetic programming.
A logical, reasonably standardized, but flexible project structure for doing and sharing data science work.
A unified framework for machine learning with time series
Open-source, low-code AutoML platform for Python. PyCaret 4.0: sklearn-native engine + React control plane.
cuDF - GPU DataFrame Library
cuDF - GPU DataFrame Library
Deep learning library featuring a higher-level API for TensorFlow.
A python library for user-friendly forecasting and anomaly detection on time series.
Automatic extraction of relevant features from time series:
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
The backtesting engine that gives you an unfair advantage. Run thousands of trading ideas before others finish one.
🧙 Build, run, and manage data pipelines for integrating and transforming data.
Out-of-Core hybrid Apache Arrow/NumPy DataFrame for Python, ML, visualization and exploration of big tabular data at a billion rows per second 🚀
Open Source Feature Flags, Experimentation, and Product Analytics
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
A roadmap connecting many of the most important concepts in machine learning, how to learn them and what tools to use to perform them.
An open source python library for automated feature engineering
H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.
Dynamic, resilient AI orchestration. Coordinate data, models, and compute as you build AI workflows.