reinforcement-learning 2 Miles Deep Dive: Main Loop, Rollout, and Reward Jun 29, 2026 GRPO from First Principles, and How verl Implements It May 19, 2026