Why the same AI model scored two different benchmark numbers
カートのアイテムが多すぎます
ご購入は五十タイトルがカートに入っている場合のみです。
カートに追加できませんでした。
しばらく経ってから再度お試しください。
ウィッシュリストに追加できませんでした。
しばらく経ってから再度お試しください。
ほしい物リストの削除に失敗しました。
しばらく経ってから再度お試しください。
ポッドキャストのフォローに失敗しました
ポッドキャストのフォロー解除に失敗しました
-
ナレーター:
-
著者:
Test a vendor's AI benchmark claim on your own work in one afternoon, before you sign a contract that rests on it.
OpenAI's o3 was announced at about 25% on a very hard math test. Months later, an independent check of the shipped model found about 10%.
- Spot the four ways an honest benchmark score gets inflated, even when nobody is lying.
- Know why one dangerous answer on your own real data is enough to stop a purchase, whatever the average says.
- Ask the one question that tells you whether a high score predicts anything about your real task.
Think you've got it? Prove it inside the program, where it goes on a record employers can check. Free to start: https://www.gage.academy/lessons/ai-governance/reproducing-a-claim-testing-a-vendor-benchmark-yourself
Where this skill is hired: AI Vendor Risk Manager, Chief Information Officer.
Episode 91 of 98 in AI Governance, the Certified AI Governance Professional (CAIGP) program from GAGE (Global Academy of Generative-AI Education). https://www.gage.academy/programs/ai-governance
adbl_web_anon_alc_button_suppression_t1
まだレビューはありません