vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Nano vLLM
Appears on
Quick read
Latest capture 2026-08-03 03:02
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +1120
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A high-throughput and memory-efficient inference and serving engine for LLMs
Minimalistic large language model 3D-parallelism training
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.