Dao-AILab/flash-attention
Fast and memory-efficient exact attention
[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Quick read
Latest capture 2026-09-04 04:07
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
96 observed captures since 2026-06-04. Observed captures are shown by default.
Stars from first capture -2
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Fast and memory-efficient exact attention
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
FlashInfer: Kernel Library for LLM Serving
High-speed Large Language Model Serving for Local Deployment
Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署.
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.