llm-d/llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
Community maintained hardware plugin for vLLM on Ascend
Appears on
Quick read
Latest capture 2026-09-11 10:54
35 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Agent workspace 28
Claude Code 3
Gemini CLI 3
11 more paths detected.
1 observed capture since 2026-09-11. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Achieve state of the art inference performance with modern accelerators on Kubernetes
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
A programmable Mixture-of-Models router for heterogeneous LLM inference
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.