lm-sys/RouteLLM
A framework for serving and evaluating LLM routers - save LLM costs without compromising quality
Switchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility - enabling flexible model selection, benchmarking, and cost/performance optimization.
Appears on
Quick read
Latest capture 2026-09-08 10:54
17 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Agent workspace 11
Claude Code 4
CodeRabbit
1 observed capture since 2026-09-08. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A framework for serving and evaluating LLM routers - save LLM costs without compromising quality
Open-source LLM router & AI cost optimizer. Routes simple prompts to cheap/local models, complex ones to premium — automatically. Drop-in OpenAI-compatible proxy for Claude Code, Codex, Cursor, OpenClaw. Saves 40-70% on AI API costs. Self-hosted, no middleman.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
LLM speculative inference server for heterogeneous hardware & consumer GPUs
A programmable Mixture-of-Models router for heterogeneous LLM inference
Open-source benchmark for browser AI agents on daily tasks.