Stanford CS329H: Machine Learning from Human Preferences | Autumn 2024 | Preference Models
Stanford Online · 79:27
This lecture builds the discrete-choice toolkit behind most modern learning-from-human-preference pipelines: observed choices are treated as noisy samples from a latent utility, and once you pick a noise model the est...