github Actively maintained

kvcache-ai/ktransformers

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

2 awesome lists

Quick read

Stars
19,150
Forks
1,503
Open issues
480
Commits
1,306

Activity and growth

Latest capture 2026-08-03 03:05

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+1955
Stored snapshots
6

Classification

Technology stack

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2024-07-26
First commit
2024-07-27
Last pushed
2026-08-02
GitHub updated
2026-08-03
Last synced
2026-08-03 03:05
Stack scanned
2026-08-03 03:05
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +1955

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

NVIDIA/Megatron-LM

Ongoing research training transformer models at scale

17,300 stars
Python 2 awesome lists

OpenNMT/CTranslate2

Fast inference engine for Transformer models

4,606 stars
C++ 1 awesome list

NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

3,371 stars
Python 2 awesome lists

InternLM/xtuner

A Next-Generation Training Engine Built for Ultra-Large MoE Models

5,170 stars
Python 2 awesome lists

ikawrakow/ik_llama.cpp

llama.cpp fork with additional SOTA quants and improved performance

2,989 stars
C++ 1 awesome list