github Actively maintained

Dao-AILab/flash-attention

Fast and memory-efficient exact attention

1 awesome list

Quick read

Stars
24,595
Forks
2,954
Open issues
1,233
Commits
1,542

Activity and growth

Latest capture 2026-08-02 03:12

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+690
Stored snapshots
6

Classification

Technology stack

Metadata

Language
Python
License
BSD-3-Clause
Default branch
main
Created
2022-05-19
First commit
2022-05-20
Last pushed
2026-08-01
GitHub updated
2026-08-02
Last synced
2026-08-02 03:12
Stack scanned
2026-08-02 03:12
Archived
No

AI development signals

2 paths

Agent instructions and tool configuration found in this repository.

Agent instructions

Claude Code

View paths

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +690

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

mit-han-lab/duo-attention

[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

539 stars
Python 0 awesome lists

certik/fastGPT

Fast GPT-2 inference written in Fortran

199 stars
Fortran 1 awesome list

EricLBuehler/candle-vllm

Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.

712 stars
Rust 1 awesome list

kvcache-ai/ktransformers

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

19,150 stars
Python 2 awesome lists

ikawrakow/ik_llama.cpp

llama.cpp fork with additional SOTA quants and improved performance

2,989 stars
C++ 1 awesome list