flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Fast and memory-efficient exact attention
Appears on
Quick read
Latest capture 2026-08-02 03:12
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +690
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
FlashInfer: Kernel Library for LLM Serving
[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Fast GPT-2 inference written in Fortran
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
llama.cpp fork with additional SOTA quants and improved performance