github Actively maintained

kaito-project/aikit

🏗️ Fine-tune, build, and deploy open-source LLMs easily!

2 awesome lists

Quick read

Stars
535
Forks
57
Open issues
35
Commits
736

Activity and growth

Latest capture 2026-08-15 03:05

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+11
Stored snapshots
6

Classification

Metadata

Language
Go
License
MIT
Default branch
main
Created
2023-09-20
First commit
2023-09-20
Last pushed
2026-08-11
GitHub updated
2026-08-11
Last synced
2026-08-15 03:05
Stack scanned
2026-08-15 03:05
Archived
No

AI development signals

3 paths

Agent instructions and tool configuration found in this repository.

Agent instructions

Claude Code

GitHub Copilot

View paths

Growth history

Tracked growth

6 observed captures since 2026-05-27. Observed captures are shown by default.

Stars from first capture +11

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

lyogavin/airllm

AirLLM 70B inference with single 4GB GPU

28,738 stars
Jupyter Notebook 3 awesome lists

ikawrakow/ik_llama.cpp

llama.cpp fork with additional SOTA quants and improved performance

2,989 stars
C++ 1 awesome list

getumbrel/llama-gpt

A self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support!

10,936 stars
TypeScript 3 awesome lists

NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

3,371 stars
Python 2 awesome lists

LostRuins/koboldcpp

Run GGUF models easily with a KoboldAI UI. One File. Zero Install.

11,312 stars
C++ 1 awesome list

defilantech/LLMKube

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.

196 stars
Go 1 awesome list