Policy Gradient in One Minute
Jia-Bin Huang · 1:18
This one-minute explainer walks the entire derivation chain of modern policy gradient methods — from the raw REINFORCE gradient up through PPO and the reward-normalization tricks used in RLHF/reasoning-model training...