Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 13: Meta RL
Stanford Online · 69:10
Meta RL trains a memory-based policy across a distribution of related MDPs so that, at test time, the agent can explore a new in-distribution task and adapt in a few episodes (or even a few timesteps) instead of learn...