NVIDIA/Megatron-LM
Ongoing research training transformer models at scale
Repo for external large-scale work
Ongoing research training transformer models at scale
A unified library of SOTA model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
Foundational Models for State-of-the-Art Speech and Text Translation
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
Fast inference engine for Transformer models
FAIR Sequence Modeling Toolkit 2
1 capture since 2026-05-27