[UCLA RL-LLM] Chapter 3.1: Reinforcement learning from human feedback (PPO, DPO)

Ernest Ryu · 45:23

RLHF is the third stage of LLM training: after pre-training and instruction fine-tuning produce a coherent, instruction-following model \(\pi\theta\), pairwise human preferences are used to train a reward model and th...

Read the full summary on tuber

Redirecting...