Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 14: Exploration

Stanford Online · 72:42

Exploration in RL is easy to analyze in bandits (via regret, UCB, posterior sampling) but is generally intractable from scratch in large MDPs such as robotics and language models; those domains need pretrained policie...

Read the full summary on tuber

Redirecting...