Open-Weight Masked Introspection: Measuring What Language Models Can Report Abou... (AI Podcast)
Daily Papers AI · 6:17
Current language models cannot reliably report whether their own internal computation was altered: across 78,000+ measurements they performed at chance (AUROC ≈ 0.5007), even though the perturbation signal was present...