sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
Appears on
Quick read
Latest capture 2026-08-03 03:11
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +383
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
SGLang is a high-performance serving framework for large language models and multimodal models.
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
A framework for few-shot evaluation of language models.
Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.
llama.cpp fork with additional SOTA quants and improved performance
A minimal Python framework for building custom AI inference servers with full control over logic, batching, and scaling.