llm-d/llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Appears on
Quick read
Latest capture 2026-08-24 03:03
1 path
Agent instructions and tool configuration found in this repository.
Agent instructions
7 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture +90
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Achieve state of the art inference performance with modern accelerators on Kubernetes
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
Kubernetes operator for deploying and managing OpenClaw AI agent instances with production-grade security, observability, and lifecycle management.
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
Connect Your Agents And Harnesses With Any Provider 🦚
Deploy any AI model, agent, database, RAG, and pipeline locally or remotely in minutes