github Actively maintained

mit-han-lab/duo-attention

[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

0 awesome lists

Quick read

Stars
539
Forks
42
Open issues
15
Commits
20

Activity and growth

Latest capture 2026-09-04 04:07

Stars · last 7 days
0 0.0%
Commits · last 7 days
0 0.0%
Stars since tracking
-2
Stored snapshots
96

Metadata

Language
Python
License
MIT
Default branch
main
Created
2024-10-15
First commit
2024-10-15
Last pushed
2025-02-10
GitHub updated
2026-08-19
Last synced
2026-09-04 04:07
Stack scanned
2026-09-04 04:07
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

96 observed captures since 2026-06-04. Observed captures are shown by default.

Stars from first capture -2

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

OptimalScale/LMFlow

An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.

8,486 stars
Python 2 awesome lists

Tiiny-AI/PowerInfer

High-speed Large Language Model Serving for Local Deployment

9,692 stars
C++ 2 awesome lists

jianzhnie/LLamaTuner

Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署.

620 stars
Python 1 awesome list

haotian-liu/LLaVA

[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.

24,981 stars
Python 1 awesome list