Dao-AILab/flash-attention
Fast and memory-efficient exact attention
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Appears on
Quick read
Latest capture 2026-09-10 10:54
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
1 observed capture since 2026-09-10. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Fast and memory-efficient exact attention
[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
FlashInfer: Kernel Library for LLM Serving
High-speed Large Language Model Serving for Local Deployment
PyTorch native quantization for training and inference
[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration