Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 5: Off-Policy Actor Critic
Stanford Online · 69:21
This lecture derives PPO and SAC from on-policy actor-critic: PPO reuses one collected batch for many policy updates by clipping importance-ratio changes so stale advantages stay usable, while SAC goes further off-pol...