Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 9: RL for LLMs
Stanford Online · 62:51
Pre-training teaches next-token prediction and a lot of world knowledge, but assistants come from a later stack: instruction fine-tuning for format, then preference optimization via RLHF or DPO so the model actually f...