xLLM-AI/xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
Appears on
Quick read
Latest capture 2026-08-02 03:07
2 paths
Agent instructions and tool configuration found in this repository.
Claude Code 2
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +162
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.
A high-throughput and memory-efficient inference and serving engine for LLMs
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
Achieve state of the art inference performance with modern accelerators on Kubernetes
Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署.