Andyyyy64/whichllm
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Hundreds of models & providers. One command to find what runs on your hardware.
Appears on
Quick read
Latest capture 2026-09-04 10:53
1 path
Agent instructions and tool configuration found in this repository.
Agent instructions
1 observed capture since 2026-09-04. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
No description.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Achieve state of the art inference performance with modern accelerators on Kubernetes
AirLLM 70B inference with single 4GB GPU
LLM training code for Databricks foundation models