Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase
AI Engineer · 21:14
Naive LLM-as-judge verifiers on web-agent benchmarks are often confidently wrong: they can score a model at 74% when a human-aligned rubric-and-screenshot verifier puts the same run at 38%, so using those judges as RL...