Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 8: Reward Learning
Stanford Online · 65:58
This lecture finishes offline RL by contrasting IQL (stay in-dataset, improve via an expectile value) with CQL (query the policy but pessimistically push down out-of-distribution Q-values), then shows why specifying r...