NVIDIA/Megatron-LM
Ongoing research training transformer models at scale
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
Appears on
Quick read
Latest capture 2026-08-03 03:08
152 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Agent workspace 126
Claude Code 24
CodeRabbit
128 more paths detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +612
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Ongoing research training transformer models at scale
Go ahead and axolotl questions
Minimalistic large language model 3D-parallelism training
AirLLM 70B inference with single 4GB GPU
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
PyTorch native quantization for training and inference