What AI gets wrong and what failure teaches us
Microsoft Research · 39:33
Jennifer Neville, a Microsoft Research partner research manager and Purdue CS/statistics professor, argues that standard single-turn benchmarks systematically overstate what users get from today’s models, and that her...