github Actively maintained

huggingface/datasets

🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools

2 awesome lists

Quick read

Stars
21,798
Forks
3,330
Open issues
1,181
Commits
4,324

Activity and growth

Latest capture 2026-08-03 03:03

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+261
Stored snapshots
6

Classification

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2020-03-26
First commit
2020-04-14
Last pushed
2026-07-31
GitHub updated
2026-08-03
Last synced
2026-08-03 03:03
Stack scanned
2026-08-03 03:03
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +261

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

src-d/datasets

source{d} datasets ("big code") for source code analysis and machine learning on source code

349 stars
Jupyter Notebook 1 awesome list

activeloopai/deeplake

Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.

9,221 stars
C++ 3 awesome lists

Agnuxo1/CAJAL

CAJAL — Local scientific paper generator. Qwen 27B fine-tuned for academic writing. AI Tribunal peer review. 6x4 CognitionBoard. 2.7x token compression. Runs offline on RTX 3090. Apache 2.0.

21 stars
Python 0 awesome lists

sfu-db/dataprep

Open-source low code data preparation library in python. Collect, clean and visualization your data in python with a few lines of code.

2,248 stars
Python 1 awesome list

Agnuxo1/p2pclaw-dataset

P2PCLAW: Training Dataset for Autonomous Scientific Peer Review - Apache 2.0

2 stars
0 awesome lists

argilla-io/distilabel

Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

3,348 stars
Python 1 awesome list