src-d/datasets
source{d} datasets ("big code") for source code analysis and machine learning on source code
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
Appears on
Quick read
Latest capture 2026-08-03 03:03
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +261
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
source{d} datasets ("big code") for source code analysis and machine learning on source code
Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.
CAJAL — Local scientific paper generator. Qwen 27B fine-tuned for academic writing. AI Tribunal peer review. 6x4 CognitionBoard. 2.7x token compression. Runs offline on RTX 3090. Apache 2.0.
Open-source low code data preparation library in python. Collect, clean and visualization your data in python with a few lines of code.
P2PCLAW: Training Dataset for Autonomous Scientific Peer Review - Apache 2.0
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.