github Actively maintained

mit-han-lab/streaming-llm

[ICLR 2024] Efficient Streaming Language Models with Attention Sinks

1 awesome list

Quick read

Stars
7,268
Forks
399
Open issues
50
Commits
35

Activity and growth

Latest capture 2026-09-07 10:55

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
0
Stored snapshots
1

Metadata

Language
Python
License
MIT
Default branch
main
Created
2023-09-29
First commit
2023-09-29
Last pushed
2024-07-11
GitHub updated
2026-09-07
Last synced
2026-09-07 10:55
Stack scanned
2026-09-07 10:55
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

1 observed capture since 2026-09-07. Observed captures are shown by default.

Stars from first capture 0

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

OptimalScale/LMFlow

An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.

8,487 stars
Python 2 awesome lists

mlabonne/llm-course

Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.

82,400 stars
3 awesome lists

vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

90,045 stars
Python 4 awesome lists

mit-han-lab/duo-attention

[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

540 stars
Python 0 awesome lists