Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Tutorial Session: Review of Q-Learning

Stanford Online · 50:39

This CS224R TA lecture by Anakiette walks from MDP notation and tabular Q-iteration on a grid world through parametric Q-learning, showing how Monte Carlo, one-step TD, and n-step returns trade bias for variance, then...

Read the full summary on tuber

Redirecting...