vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Appears on
Quick read
Latest capture 2026-08-03 03:04
7 paths
Agent instructions and tool configuration found in this repository.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +112
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
Achieve state of the art inference performance with modern accelerators on Kubernetes
Universal LLM Deployment Engine with ML Compilation
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.