github Actively maintained

NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

2 awesome lists

Quick read

Stars
3,371
Forks
524
Open issues
325
Commits
1,089

Activity and growth

Latest capture 2026-08-03 03:08

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+612
Stored snapshots
6

Classification

Technology stack

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2024-04-23
First commit
2024-05-08
Last pushed
2026-08-03
GitHub updated
2026-08-03
Last synced
2026-08-03 03:08
Stack scanned
2026-08-03 03:08
Archived
No

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +612

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

NVIDIA/Megatron-LM

Ongoing research training transformer models at scale

17,300 stars
Python 2 awesome lists

huggingface/nanotron

Minimalistic large language model 3D-parallelism training

2,771 stars
Python 2 awesome lists

lyogavin/airllm

AirLLM 70B inference with single 4GB GPU

28,738 stars
Jupyter Notebook 3 awesome lists

kaito-project/aikit

🏗️ Fine-tune, build, and deploy open-source LLMs easily!

535 stars
Go 2 awesome lists

pytorch/ao

PyTorch native quantization for training and inference

2,962 stars
Python 1 awesome list