John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI

Dwarkesh Patel · 95:50

Pre-training produces a calibrated next-token imitator of the web that can take on many personas; post-training then narrows that into a helpful assistant optimized for human-useful outputs rather than raw imitation....

Read the full summary on tuber

Redirecting...