lyogavin/airllm
AirLLM 70B inference with single 4GB GPU
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
Appears on
Quick read
Latest capture 2026-08-15 03:05
3 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Claude Code
GitHub Copilot
6 observed captures since 2026-05-27. Observed captures are shown by default.
Stars from first capture +11
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
AirLLM 70B inference with single 4GB GPU
llama.cpp fork with additional SOTA quants and improved performance
A self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support!
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
Run GGUF models easily with a KoboldAI UI. One File. Zero Install.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.