llm-d/llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
A Datacenter Scale Distributed Inference Serving Framework
Appears on
Quick read
Latest capture 2026-08-02 03:07
176 paths
Agent instructions and tool configuration found in this repository.
Agent instructions 29
Agent workspace 117
Claude Code 29
CodeRabbit
152 more paths detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +579
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Achieve state of the art inference performance with modern accelerators on Kubernetes
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
Low-code framework for building custom LLMs, neural networks, and other AI models
Build, run and scale AI agents like API and microservices
A blazing fast AI Gateway with integrated guardrails. Route to 1,600+ LLMs, 50+ AI Guardrails with 1 fast & friendly API.
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.