EricLBuehler/candle-vllm
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
LLM speculative inference server for heterogeneous hardware & consumer GPUs
Appears on
Quick read
Latest capture 2026-09-07 10:54
1 path
Agent instructions and tool configuration found in this repository.
Codex
1 observed capture since 2026-09-07. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
Low-code framework for building custom LLMs, neural networks, and other AI models
Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
Fast, flexible LLM inference
Achieve state of the art inference performance with modern accelerators on Kubernetes