Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Unified MySQL, Postgres & FlightSQL Server, Powered by DuckDB.
Obsidian tools - a Python package for analysing an Obsidian.md vault
Engine for AI/ML/Data tracking, visualization, explainability, drift detection, and dashboards for Polyaxon.
Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.
A scikit-learn-compatible Python implementation of ReBATE, a suite of Relief-based feature selection algorithms for Machine Learning.
Datasets & Analyses for Formula 1 World Championship
Core meta for awesome-public-datasets. Contribute new data here!
Spatial Representations for Artificial Intelligence - a Python library toolkit for geospatial machine learning focused on creating embeddings for downstream tasks
A Python 3 library making time series data mining tasks, utilizing matrix profile algorithms, accessible to everyone.
Fast DuckDB-powered TypeScript library for tabular, geospatial, vector, AI, Google Sheets, and data visualization workflows on Deno, Node.js, and Bun.
A unified framework for tabular probabilistic regression, time-to-event prediction, and probability distributions in python
The Machine Learning project including ML/DL projects, notebooks, cheat codes of ML/DL, useful information on AI/AGI and codes or snippets/scripts/tasks with tips.
A pure Python implementation of Apache Spark's RDD and DStream interfaces.
Python-based implementations of algorithms for learning on imbalanced data.
A fast xgboost feature selection algorithm
A collection of Jupyter Notebooks for learning Python for Data Science.
Explore and analyse your Home Assistant data
BatchFlow helps you conveniently work with random or sequential batches of your data and define data processing and machine learning workflows even for datasets that do not fit into memory.
AI-powered NBA game outcome predictor that uses advanced team stats and trend-based features to forecast winners and track model performance
CLI utility and Python module for analyzing log files and other data.
Proton Exchange Membrane (PEM) Fuel Cell Dataset
Python package implementing ML feature engineering and pre-processing for polars or pandas dataframes.
AI-powered SDK featuring algorithmic trading, backtesting, deployment on 100+ exchanges, and multiple optimization engines.
Fast, accurate, open-source geocoding in Python
Data Analysis With Vibes
PSDuckDB is a PowerShell module that provides seamless integration with DuckDB, enabling efficient execution of analytical SQL queries directly from the PowerShell environment.
Build, test, deploy, iterate - Dev and prod tool for data science pipelines
A statistical computing toolkit for DuckDB.
An open and opinionated trading platform using productive & familiar open source libraries and tools for strategy research, execution and operation.
Common Lisp CFFI wrapper around the DuckDB C API