Tiiny-AI/PowerInfer
High-speed Large Language Model Serving for Local Deployment
FlashInfer: Kernel Library for LLM Serving
Appears on
Quick read
Latest capture 2026-08-03 03:02
10 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Claude Code 9
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +412
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
High-speed Large Language Model Serving for Local Deployment
Fast and memory-efficient exact attention
A high-throughput and memory-efficient inference and serving engine for LLMs
[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.