Stanford CS221 | Autumn 2025 | Lecture 9: Policy Gradient

Stanford Online · 74:02

This lecture is the second in a reinforcement learning series: after tabular Q-learning and SARSA, it shows how to scale to huge state spaces with function approximation and how to optimize a policy directly via polic...

Read the full summary on tuber

Redirecting...