github Actively maintained

thu-ml/SageAttention

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

1 awesome list

Quick read

Stars
3,717
Forks
504
Open issues
213
Commits
188

Activity and growth

Latest capture 2026-09-10 10:54

Stars · last 7 days
0 0.0%
Commits · last 7 days
0 0.0%
Stars since tracking
0
Stored snapshots
1

Classification

Metadata

Language
Cuda
License
Apache-2.0
Default branch
main
Created
2024-10-03
First commit
2024-10-03
Last pushed
2026-01-17
GitHub updated
2026-09-10
Last synced
2026-09-10 10:54
Stack scanned
2026-09-10 10:54
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

1 observed capture since 2026-09-10. Observed captures are shown by default.

Stars from first capture 0

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

mit-han-lab/duo-attention

[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

540 stars
Python 0 awesome lists

Tiiny-AI/PowerInfer

High-speed Large Language Model Serving for Local Deployment

9,787 stars
C++ 2 awesome lists

pytorch/ao

PyTorch native quantization for training and inference

2,977 stars
Python 1 awesome list

mit-han-lab/llm-awq

[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

3,631 stars
Python 1 awesome list