How LLMs Learn to Reason [GRPO]
Jia-Bin Huang · 23:31
This video traces the full lineage of RL algorithms behind LLM reasoning — from vanilla policy gradients through actor-critic, TRPO, PPO, GRPO, and finally Dr. GRPO — showing how each step fixes a concrete flaw in the...