Dao-AILab/flash-attention
Fast and memory-efficient exact attention
Repository profile
[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Repository updates
Get generated mit-han-lab/duo-attention development summaries by email, or follow the weekly and monthly RSS feeds.
Sign in to subscribe by email. RSS feeds are public.
Sign in to subscribeTracked growth, recent movement, and commit velocity from stored repository snapshots.
Latest capture 2026-07-20 04:06
50 captures since 2026-06-04
Stars from baseline -2
All tracked data
Frameworks, package managers, ecosystems, and dependency manifests found during catalog scans.
Scanned 2026-07-20 04:06
Searchable topics, generated tags, and stack labels that explain where this repository fits.
Agent instructions and tool configuration paths found in the repository tree.
Nearest indexed repositories by embedding similarity.
Fast and memory-efficient exact attention
FlashInfer: Kernel Library for LLM Serving
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
High-speed Large Language Model Serving for Local Deployment
Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署.
PyTorch native quantization and sparsity for training and inference