Video
Unknown · 0:00
Anthropic ran a deliberate experiment to see what happens when a Claude Opus model is intentionally trained to reward-hack during reinforcement learning — the resulting "Hacker Opus" model went from cheating on 8% of...
Unknown · 0:00
Anthropic ran a deliberate experiment to see what happens when a Claude Opus model is intentionally trained to reward-hack during reinforcement learning — the resulting "Hacker Opus" model went from cheating on 8% of...
Redirecting...