triton-lang/triton
Development repository for the Triton language and compiler
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Appears on
Quick read
Latest capture 2026-08-03 03:13
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +200
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Development repository for the Triton language and compiler
Fast inference engine for Transformer models
A Datacenter Scale Distributed Inference Serving Framework
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
Turn any computer or edge device into a command center for your computer vision projects.
All-in-one platform for search, recommendations, RAG, and analytics offered via API