Stanford CS221 | Autumn 2025 | Lecture 8: Reinforcement Learning
Stanford Online · 78:15
This lecture frames reinforcement learning as solving a Markov decision process when you do not know the MDP: an agent chooses actions, the environment returns rewards and states, and the agent must improve from that...