github Actively maintained

huggingface/tokenizers

馃挜 Fast State-of-the-Art Tokenizers optimized for Research and Production

Quick read

Stars
10,943
Forks
1,161
Open issues
263
Commits
2,004

Activity and growth

Latest capture 2026-08-03 03:04

Stars 路 last 7 days
No history
Commits 路 last 7 days
No history
Stars since tracking
+175
Stored snapshots
6

Classification

Metadata

Language
Rust
License
Apache-2.0
Default branch
main
Created
2019-11-01
First commit
2019-11-01
Last pushed
2026-08-03
GitHub updated
2026-08-02
Last synced
2026-08-03 03:04
Stack scanned
2026-08-03 03:04
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +175

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

openai/tiktoken

tiktoken is a fast BPE tokeniser for use with OpenAI's models.

18,970 stars
Python 1 awesome list

pytorch/text

Models, data loaders and abstractions for language processing, powered by PyTorch

3,554 stars
Python 1 awesome list

botisan-ai/gpt3-tokenizer

Isomorphic JavaScript/TypeScript Tokenizer for GPT-3 and Codex Models by OpenAI.

170 stars
TypeScript 1 awesome list

huggingface/transformers

馃 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

163,267 stars
Python 6 awesome lists

huggingface/peft

馃 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.

21,547 stars
Python 3 awesome lists