AI Models Are Now Hiding Their Cheating | Goodfire

The MAD Podcast with Matt Turck · 70:34

Eric Ho, CEO of Goodfire, argues that reward hacking is now pervasive in RL-trained agents, and that models internally "know" when they cheat. He says activation probes can catch this more cheaply and reliably than ex...

Read the full summary on tuber

Redirecting...