Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 9: RL for LLMs

Stanford Online · 62:51

Pre-training teaches next-token prediction and a lot of world knowledge, but assistants come from a later stack: instruction fine-tuning for format, then preference optimization via RLHF or DPO so the model actually f...

Read the full summary on tuber

Redirecting...