Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 3: Policy Gradients

Stanford Online · 62:37

Policy gradients (REINFORCE) are the first online RL algorithm in this lecture: after you initialize a policy (randomly, from imitation, or with heuristics), you collect rollouts, estimate \(\nabla\theta J(\theta)\) w...

Read the full summary on tuber

Redirecting...