Reinforcement learning · Courses & experiments
Hands-on Modern RL
Modern reinforcement learning: from fundamentals to LLMs and agents
Combine concepts, mathematical derivations and coding experiments to learn value functions, DQN, policy gradients and PPO, then explore RLHF, DPO, GRPO and Agentic RL.
Start hereRead the introduction and setup instructions, then complete the chapter experiments. The course is still developing; start with completed chapters.

