OptimalScale/LMFlow
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
[ICLR 2024] Efficient Streaming Language Models with Attention Sinks
Appears on
Quick read
Latest capture 2026-09-07 10:55
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
1 observed capture since 2026-09-07. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
DeepSeek LLM: Let there be answers
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
A high-throughput and memory-efficient inference and serving engine for LLMs
Code and documents of LongLoRA and LongAlpaca (ICLR 2024 Oral)
[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads