xLLM-AI/xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.
Appears on
Quick read
Latest capture 2026-06-28 03:08
34 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Agent workspace 27
Claude Code 3
Gemini CLI 3
10 more paths detected.
4 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +70
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
A high-throughput and memory-efficient inference and serving engine for LLMs
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
Achieve state of the art inference performance with modern accelerators on Kubernetes
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Universal LLM Deployment Engine with ML Compilation