Data-Centric-AI-Community/fg-data-profiling
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
Repository profile
A unified interface for distributed computing. Fugue executes SQL, Python, Pandas, and Polars code on Spark, Dask and Ray without any rewrites.
Repository updates
Get generated fugue-project/fugue development summaries by email, or follow the weekly and monthly RSS feeds.
Sign in to subscribe by email. RSS feeds are public.
Sign in to subscribeTracked growth, recent movement, and commit velocity from stored repository snapshots.
Latest capture 2026-08-02 13:54
1 capture since 2026-08-02
Stars from baseline 0
All tracked data
Frameworks, package managers, ecosystems, and dependency manifests found during catalog scans.
Scanned 2026-08-02 13:54
pyproject.toml
python ecosystem,
44 dependencies
uv.lock
python ecosystem,
0 dependencies
Searchable topics, generated tags, and stack labels that explain where this repository fits.
Agent instructions and tool configuration paths found in the repository tree.
Nearest indexed repositories by embedding similarity.
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
Apache DataFusion SQL Query Engine
Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.
Simple and Distributed Machine Learning
🦀 event stream processing for developers to collect and transform data in motion to power responsive data intensive applications.
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.