Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Meltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.
Malloy is a modern open source language for describing data relationships and transformations.
Data-centric LLM training with dynamic sample selection, domain mixture optimization, and example reweighting inside the LLaMA-Factory training loop.
ArcticDB is a high performance, serverless DataFrame database built for the Python Data Science ecosystem.
LLM based data scientist, AI native data application. AI-driven infinite thinking redefines BI.
AI code-writing assistant that understands data content
A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.
The modern replacement for Jupyter Notebooks
A lightweight opinionated ETL framework, halfway between plain scripts and Apache Airflow
An Internet-Scale Database.
In-memory tabular data in Julia
Interactively find and recover deleted or :point_right: overwritten :point_left: files from your terminal
Powerful Analytics Solution. Setup in 30 seconds. Display all your data on a Simple, AI-powered dashboard. Fully self-hostable and GDPR compliant. Alternative to Google Analytics, MixPanel, Plausible, Umami & Matomo.
Draw datasets from within Python notebooks.
👾 nao is an open source analytics agent. (1) Create context with nao-core cli, (2) deploy nao chat interface for everyone
Machine learning with dataframes
a standard filetree for /r/datacurator [ and r/datahoarder ]
Concurrent Python made simple
Clean APIs for data cleaning. Python implementation of R package Janitor
visual data prep powered by python
A library for building rich, web-based geospatial 2D & 3D data platforms.
Visualize and share your data. All in SQL. Powered by DuckDB.
SQL-based streaming analytics platform at scale
Epsilla is a high performance Vector Database Management System
Easy pipelines for pandas DataFrames.
This repository has been merged into Canner/WrenAI under the core/ directory
Pandas, Polars, Spark, and Snowpark DataFrame comparison for humans and more!
Python SDK for IEX Cloud
Port of a popular ruby faker gem written in kotlin. Generate realistically looking fake data such as names, addresses, banking details, and many more, that can be used for testing and data anonymization purposes.
How to disable Firefox Telemetry and Data Collection