Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 3: Policy Gradients
Stanford Online · 62:37
Policy gradients (REINFORCE) are the first online RL algorithm in this lecture: after you initialize a policy (randomly, from imitation, or with heuristics), you collect rollouts, estimate \(\nabla\theta J(\theta)\) w...