llm-d/llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
A minimal Python framework for building custom AI inference servers with full control over logic, batching, and scaling.
Appears on
Quick read
Latest capture 2026-08-03 03:05
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +39
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Achieve state of the art inference performance with modern accelerators on Kubernetes
Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Integrates with 50+ LLM Providers, VectorDBs, Agent Frameworks and GPUs.
A high-throughput and memory-efficient inference and serving engine for LLMs
Low-code framework for building custom LLMs, neural networks, and other AI models
AirLLM 70B inference with single 4GB GPU
Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.