AI Interpretability, Safety, and Meaning - Nora Belrose
Machine Learning Street Talk · 149:50
Nora Belrose (Head of Interpretability, EleutherAI) walks through her technical work on linear concept erasure (LEACE/QLEACE) and the "simplicity bias" of neural networks, then spends the bulk of the interview on phil...