OptimalScale/LMFlow
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
High-speed Large Language Model Serving for Local Deployment
Appears on
Quick read
Latest capture 2026-08-03 03:13
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +198
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
FlashInfer: Kernel Library for LLM Serving
Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署.
AirLLM 70B inference with single 4GB GPU
SGLang is a high-performance serving framework for large language models and multimodal models.
QLoRA: Efficient Finetuning of Quantized LLMs