ikawrakow/ik_llama.cpp
llama.cpp fork with additional SOTA quants and improved performance
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
Appears on
Quick read
Latest capture 2026-08-15 03:05
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-27. Observed captures are shown by default.
Stars from first capture +45
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
llama.cpp fork with additional SOTA quants and improved performance
Fast, flexible LLM inference
Minimalist ML framework for Rust
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
LLM inference in C/C++