Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Learn to build your Second Brain AI assistant with LLMs, agents, RAG, fine-tuning, LLMOps and AI systems techniques.
Meltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.
Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.
A new SOTA for RAG — an original retrieval architecture and an open-source knowledge base for humans and agents.
A low code Machine Learning personalized ranking service for articles, listings, search results, recommendations that boosts user engagement. A friendly Learn-to-Rank engine
Data Contracts engine for the modern data stack. https://www.soda.io
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
👾 nao is an open source analytics agent. (1) Create context with nao-core cli, (2) deploy nao chat interface for everyone
Concurrent Python made simple
Clean APIs for data cleaning. Python implementation of R package Janitor
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.
Polyglot workflows without leaving the comfort of your technology stack.
Synmetrix – production-ready open source semantic layer on Cube
Browser-based data analytics & BI with DuckDB-WASM. Privacy-first: your data stays local; no server required.
A MCP (Model Context Protocol) server for interacting with dbt.
Unified MySQL, Postgres & FlightSQL Server, Powered by DuckDB.
Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.
The universal metrics layer. Compatible with 15+ formats: Cube, MetricFlow, LookML, Omni, BSL, LDM, Cortex, Malloy, OSI, SML, TML, Hex, Rill, Superset
A DuckDB extension to read data directly from databases supporting the ODBC interface
ETL / ELT / Reverse ETL Framework powered by DuckDB, designed to seamlessly integrate and process data from diverse sources. It leverages Markdown as a configuration medium, where YAML blocks define metadata for each data source, and embedded SQL blocks specify the extraction, transformation, and loading logic.
A lightweight Snowflake emulator built with Go and DuckDB for local development and testing
Lightning-fast validation in python.
Convert a SQLite database into a DuckDB database — tables, indexes and views, in one command
Staff-level engineering judgment for coding agents, from design to production.
Your data, explored locally — and your AI agents kept on a leash. A federated data explorer with governed agentic access.
Virtual warehouse — SQL + Jupyter + AI over cloud Parquet via DuckDB