Why High Benchmark Scores Don’t Mean Better AI [SPONSORED]

Machine Learning Street Talk · 16:05

Andrew Gordon and Nora Petrova of Prolific argue that the LLM evaluation field over-relies on technical benchmarks (MMLU, Humanity's Last Exam) that say nothing about whether a model is pleasant or safe to actually us...

Read the full summary on tuber

Redirecting...