Stanford CS329H: Machine Learning from Human Preferences | Autumn 2024 | Preference Models

Stanford Online · 79:27

This lecture builds the discrete-choice toolkit behind most modern learning-from-human-preference pipelines: observed choices are treated as noisy samples from a latent utility, and once you pick a noise model the est...

Read the full summary on tuber

Redirecting...