llm-d/llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
A high-throughput and memory-efficient inference and serving engine for LLMs
Quick read
Latest capture 2026-08-26 03:03
28 paths
Agent instructions and tool configuration found in this repository.
Agent instructions 3
Agent workspace 13
Claude Code 9
Codex
Gemini CLI 2
4 more paths detected.
8 observed captures since 2026-05-22. Observed captures are shown by default.
Stars from first capture +9299
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Achieve state of the art inference performance with modern accelerators on Kubernetes
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
SGLang is a high-performance serving framework for large language models and multimodal models.