[UCLA RL-LLM] Chapter 1.2: Deep policy evaluation

Ernest Ryu · 27:06

Policy evaluation asks how good a given policy π is (via its value function), and this lecture shows how to approximate that with Monte Carlo rollouts versus temporal-difference (TD) bootstrapping—first for V, then fo...

Read the full summary on tuber

Redirecting...