[UCLA RL-LLM] Chapter 3.1: Reinforcement learning from human feedback (PPO, DPO)
Ernest Ryu · 45:23
RLHF is the third stage of LLM training: after pre-training and instruction fine-tuning produce a coherent, instruction-following model \(\pi\theta\), pairwise human preferences are used to train a reward model and th...