Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 10: RL for LLM Reasoning
Stanford Online · 70:30
Next-token supervised training on human math solutions cannot reach expert-level reasoning because high-quality traces are scarce, so models produce exam-looking writeups that are logically wrong; RL helps by treating...