What Your AI Benchmark is Really Telling You
カートのアイテムが多すぎます
カートに追加できませんでした。
ウィッシュリストに追加できませんでした。
ほしい物リストの削除に失敗しました。
ポッドキャストのフォローに失敗しました
ポッドキャストのフォロー解除に失敗しました
-
ナレーター:
-
著者:
Learn more about the host Laurence Gill at www.laurencegill.com.
In Episode 12, Laurence Gill takes that number apart. A Stanford research team called BetterBench built a 46-point audit covering benchmark design, reproducibility, and documentation, then scored 24 widely-cited tests against it. MMLU came in at 5.5. GPQA, a far less publicized test, scored double that. The reasons are specific: ambiguous question phrasing that swings scores when a comma moves, a reproducibility gap across most published benchmarks, and a quiet contamination problem where models may have already seen the answer key buried somewhere in their training data.
Laurence walks through how these tests actually work, why Goodhart’s Law explains the industry’s race to game them, and how newer benchmarks like GPQA and ARC-AGI are trying to close the gap. It closes with five questions to run through before any benchmark score is allowed to inform a real decision and one open question about what happens when AI starts writing the tests that grade other AI.