Stanford CS221 | Autumn 2025 | Lecture 9: Policy Gradient
Stanford Online · 74:02
This lecture is the second in a reinforcement learning series: after tabular Q-learning and SARSA, it shows how to scale to huge state spaces with function approximation and how to optimize a policy directly via polic...