Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 13: Meta RL

Stanford Online · 69:10

Meta RL trains a memory-based policy across a distribution of related MDPs so that, at test time, the agent can explore a new in-distribution task and adapt in a few episodes (or even a few timesteps) instead of learn...

Read the full summary on tuber

Redirecting...