microsoft/SynapseML
Simple and Distributed Machine Learning
Repository profile
Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends
Repository updates
Get generated moj-analytical-services/splink development summaries by email, or follow the weekly and monthly RSS feeds.
Sign in to subscribe by email. RSS feeds are public.
Sign in to subscribeTracked growth, recent movement, and commit velocity from stored repository snapshots.
Latest capture 2026-08-02 13:56
1 capture since 2026-08-02
Stars from baseline 0
All tracked data
Frameworks, package managers, ecosystems, and dependency manifests found during catalog scans.
Scanned 2026-08-02 13:56
pyproject.toml
python ecosystem,
32 dependencies
uv.lock
python ecosystem,
0 dependencies
docs/demos/pyproject.toml
python ecosystem,
10 dependencies
docs/demos/uv.lock
python ecosystem,
0 dependencies
Searchable topics, generated tags, and stack labels that explain where this repository fits.
Agent instructions and tool configuration paths found in the repository tree.
Nearest indexed repositories by embedding similarity.
Simple and Distributed Machine Learning
PySpark + Scikit-learn = Sparkit-learn
Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.
A unified interface for distributed computing. Fugue executes SQL, Python, Pandas, and Polars code on Spark, Dask and Ray without any rewrites.
Open-source low code data preparation library in python. Collect, clean and visualization your data in python with a few lines of code.
Project SnappyData - memory optimized analytics database, based on Apache Spark™ and Apache Geode™. Stream, Transact, Analyze, Predict in one cluster