github Actively maintained

vllm-project/llm-compressor

Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

2 awesome lists

Quick read

Stars
3,770
Forks
657
Open issues
126
Commits
3,234

Activity and growth

Latest capture 2026-09-10 10:55

Stars · last 7 days
0 0.0%
Commits · last 7 days
0 0.0%
Stars since tracking
0
Stored snapshots
1

Classification

Technology stack

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2024-06-20
First commit
2019-07-26
Last pushed
2026-09-10
GitHub updated
2026-09-10
Last synced
2026-09-10 10:55
Stack scanned
2026-09-10 10:55
Archived
No

Growth history

Tracked growth

1 observed capture since 2026-09-10. Observed captures are shown by default.

Stars from first capture 0

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

mit-han-lab/llm-awq

[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

3,631 stars
Python 1 awesome list

NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

3,780 stars
Python 2 awesome lists

intel/neural-compressor

SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime

2,707 stars
Python 1 awesome list

vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

90,045 stars
Python 4 awesome lists

InternLM/lmdeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

8,050 stars
Python 3 awesome lists