Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 8: Reward Learning

Stanford Online · 65:58

This lecture finishes offline RL by contrasting IQL (stay in-dataset, improve via an expectile value) with CQL (query the policy but pessimistically push down out-of-distribution Q-values), then shows why specifying r...

Read the full summary on tuber

Redirecting...