『What Your AI Benchmark is Really Telling You』のカバーアート

What Your AI Benchmark is Really Telling You

What Your AI Benchmark is Really Telling You

無料で聴く

ポッドキャストの詳細を見る

Learn more about the host Laurence Gill at www.laurencegill.com.

In Episode 12, Laurence Gill takes that number apart. A Stanford research team called BetterBench built a 46-point audit covering benchmark design, reproducibility, and documentation, then scored 24 widely-cited tests against it. MMLU came in at 5.5. GPQA, a far less publicized test, scored double that. The reasons are specific: ambiguous question phrasing that swings scores when a comma moves, a reproducibility gap across most published benchmarks, and a quiet contamination problem where models may have already seen the answer key buried somewhere in their training data.


Laurence walks through how these tests actually work, why Goodhart’s Law explains the industry’s race to game them, and how newer benchmarks like GPQA and ARC-AGI are trying to close the gap. It closes with five questions to run through before any benchmark score is allowed to inform a real decision and one open question about what happens when AI starts writing the tests that grade other AI.

adbl_web_anon_alc_button_suppression_t1
まだレビューはありません