lyogavin/airllm
AirLLM 70B inference with single 4GB GPU
Official inference framework for 1-bit LLMs
Appears on
Quick read
Latest capture 2026-08-03 03:05
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +701
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
AirLLM 70B inference with single 4GB GPU
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
llama.cpp fork with additional SOTA quants and improved performance
High-speed Large Language Model Serving for Local Deployment
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
Fast Multimodal LLM on Mobile Devices