Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 14: Exploration
Stanford Online · 72:42
Exploration in RL is easy to analyze in bandits (via regret, UCB, posterior sampling) but is generally intractable from scratch in large MDPs such as robotics and language models; those domains need pretrained policie...