Open-Weight Masked Introspection: Measuring What Language Models Can Report Abou... (AI Podcast)

Daily Papers AI · 6:17

Current language models cannot reliably report whether their own internal computation was altered: across 78,000+ measurements they performed at chance (AUROC ≈ 0.5007), even though the perturbation signal was present...

Read the full summary on tuber

Redirecting...