Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.
Incremental engine for long horizon agents 🌟 Star if you like it!
Miller is like awk, sed, cut, join, and sort for name-indexed data such as CSV, TSV, and tabular JSON
High-performance AI pipeline engine with a C++ core and 50+ Python-extensible nodes. Build, debug, and scale LLM workflows with 13+ model providers, 8+ vector databases, and agent orchestration, all from your IDE. Includes VS Code extension, TypeScript/Python SDKs, and Docker deployment.
Unified querying, transformation, and modification of JSON, TOML, YAML, XML, INI, HCL, KDL and CSV.
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
A lightweight data processing framework built on DuckDB and 3FS.
A light-weight, flexible, and expressive statistical data testing library
Kubernetes-native platform to run massively parallel data/streaming jobs
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
Concurrent Python made simple
A pure Python implementation of Apache Spark's RDD and DStream interfaces.
Interactive F1 strategy dashboard tutorial using the OpenF1 API, Streamlit, Plotly, and pandas.