gotzmann/llama.go
llama.go is like llama.cpp in pure Golang!
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
Quick read
Latest capture 2026-08-03 03:07
3 paths
Agent instructions and tool configuration found in this repository.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +988
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
llama.go is like llama.cpp in pure Golang!
LLM inference in C/C++
Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
A self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support!
Utilities intended for use with Llama models.