Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
Dwarkesh Patel · 144:01
RL with clean, verifiable rewards has finally proven it can produce expert-level performance, and Sholto Douglas and Trenton Bricken (both at Anthropic) argue that competent, independent software-engineering agents sh...