github Actively maintained

vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

Quick read

Stars
90,045
Forks
21,204
Open issues
7,012
Commits
20,408

Activity and growth

Latest capture 2026-08-26 03:03

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+9299
Stored snapshots
8

Classification

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2023-02-09
First commit
2023-02-09
Last pushed
2026-08-26
GitHub updated
2026-08-26
Last synced
2026-08-26 03:03
Stack scanned
2026-08-26 03:03
Archived
No

Growth history

Tracked growth

8 observed captures since 2026-05-22. Observed captures are shown by default.

Stars from first capture +9299

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

llm-d/llm-d

Achieve state of the art inference performance with modern accelerators on Kubernetes

3,959 stars
Shell 2 awesome lists

xLLM-AI/xllm

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

1,555 stars
C++ 0 awesome lists

jd-opensource/xllm

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.

1,370 stars
C++ 1 awesome list

InternLM/lmdeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

7,985 stars
Python 3 awesome lists

alibaba/rtp-llm

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

1,292 stars
Cuda 1 awesome list

sgl-project/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

31,109 stars
Python 5 awesome lists