llm-d/llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
SGLang is a high-performance serving framework for large language models and multimodal models.
Quick read
Latest capture 2026-08-03 03:12
139 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Claude Code 138
115 more paths detected.
8 observed captures since 2026-05-22. Observed captures are shown by default.
Stars from first capture +2983
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Achieve state of the art inference performance with modern accelerators on Kubernetes
A framework for few-shot evaluation of language models.
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
No description.
llama.cpp fork with additional SOTA quants and improved performance
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.