SkyRL: Democratizing RL Training for Researchers
Agents have taken center stage in 2026, and as they shift to longer-horizon, multi-turn interactions, the systems challenges and requirements on training infrastructure have evolved with them.
Sumanth Hegde and Eric Tang from Anyscale trace SkyRL from a set of modular APIs decoupling training, inference, and environments to where it is today: scalable, fully async RL training on 350B+ parameter MoE models with Megatron and vLLM, the multi-tenant Tinker Engine that lets researchers use their own hardware with Tinker's flexible training APIs, a redesign toward HTTP-based APIs for scalable inference with native RL API contributions to vLLM, and community recipes including large MoE training on long-horizon tasks and custom recursive language models.
You'll leave with a picture of modular, performant RL training and where SkyRL is headed.