Stanford CS329A Self-Improving AI Agents | Part 8 | Agentic Evaluations and Long Horizon Tasks
Stanford Online · 75:17
Traditional chatbot and single-shot QA benchmarks are saturating, so this lecture argues that agentic progress has to be measured on three complementary axes—how long a task a model can finish (METR time horizons), wh...