EleutherAI/gpt-neox
An implementation of model parallel autoregressive transformers on GPUs, based on the Megatron and DeepSpeed libraries
Scalable toolkit for efficient model reinforcement
Appears on
Quick read
Latest capture 2026-08-03 03:06
40 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Agent workspace 21
Claude Code 18
16 more paths detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +214
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
An implementation of model parallel autoregressive transformers on GPUs, based on the Megatron and DeepSpeed libraries
Minimalistic large language model 3D-parallelism training
A framework for few-shot evaluation of language models.
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
Ongoing research training transformer models at scale