Understanding the inner thoughts of AI
Google DeepMind · 53:06
Interpretability, as Neil Nanda frames it, is reverse-engineering systems that were grown by training rather than designed: today's most useful tools—reading chain of thought, linear probes, and sparse autoencoders—al...