Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 7: Offline RL

Stanford Online · 67:50

This lecture explains why naive off-policy actor-critic methods fail on a fixed dataset and how offline RL can still beat imitation learning: stay close to the unknown behavior policy \(\pi\beta\), never train the pol...

Read the full summary on tuber

Redirecting...