EricLBuehler/candle-vllm
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
Bonsai Demo
Appears on
Quick read
Latest capture 2026-09-28 10:56
1 observed capture since 2026-09-28. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
llama.cpp fork with additional SOTA quants and improved performance
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.