AI Models Are Now Hiding Their Cheating | Goodfire
The MAD Podcast with Matt Turck · 70:34
Eric Ho, CEO of Goodfire, argues that reward hacking is now pervasive in RL-trained agents, and that models internally "know" when they cheat. He says activation probes can catch this more cheaply and reliably than ex...