github Actively maintained

flashinfer-ai/flashinfer

FlashInfer: Kernel Library for LLM Serving

2 awesome lists

Quick read

Stars
6,086
Forks
1,223
Open issues
894
Commits
2,680

Activity and growth

Latest capture 2026-08-03 03:02

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+412
Stored snapshots
6

Classification

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2023-07-22
First commit
2023-07-22
Last pushed
2026-08-03
GitHub updated
2026-08-03
Last synced
2026-08-03 03:02
Stack scanned
2026-08-03 03:02
Archived
No

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +412

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

Tiiny-AI/PowerInfer

High-speed Large Language Model Serving for Local Deployment

9,692 stars
C++ 2 awesome lists

vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

90,045 stars
Python 4 awesome lists

mit-han-lab/duo-attention

[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

539 stars
Python 0 awesome lists

xLLM-AI/xllm

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

1,555 stars
C++ 0 awesome lists

OptimalScale/LMFlow

An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.

8,486 stars
Python 2 awesome lists