xLLM-AI/xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
Universal LLM Deployment Engine with ML Compilation
Appears on
Quick read
Latest capture 2026-08-03 03:09
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +311
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.
MLX: An array framework for Apple silicon
Achieve state of the art inference performance with modern accelerators on Kubernetes