My Instruments Lied to Me Six Times
Building and Testing LLM Systems in Production, and How to Find Out What Is Actually True
カートのアイテムが多すぎます
ご購入は五十タイトルがカートに入っている場合のみです。
カートに追加できませんでした。
しばらく経ってから再度お試しください。
ウィッシュリストに追加できませんでした。
しばらく経ってから再度お試しください。
ほしい物リストの削除に失敗しました。
しばらく経ってから再度お試しください。
ポッドキャストのフォローに失敗しました
ポッドキャストのフォロー解除に失敗しました
Audibleプレミアムプラン30日間無料体験
オーディオブック・ポッドキャスト・オリジナル作品など数十万以上の対象作品が聴き放題。
オーディオブックをお得な会員価格で購入できます。
30日間の無料体験後は月額¥1500で自動更新します。いつでも退会できます。
¥2,010 で購入
-
ナレーター:
-
AI Voice A synthetic voice
-
著者:
-
Nikolaos Broikos
この作品は、デジタルボイスによる朗読を使用しています。
デジタルボイスは、オーディオブック用にコンピューター生成された朗読です。
One of them was the strongest finding in the entire project: a perfect result, every trial, in the direction the whole field expects. It was completely false, and the only thing that caught it was opening a file and reading what the model had actually said.
This is the instrument book. How to build a comparison that means something. How to find out what a wrong answer scores on your own tests, before you report a score at all. How to prove a grader can fail before you let it grade anything. Why an evaluator's error rate tells you almost nothing and its error direction tells you everything.
Every method in it is a few hundred lines of ordinary code with no dependencies, and every one of them exists because something went wrong first.
It is also, unusually, a book that documents its own failures at full length: the detector that looked for one set form of words and reported ninety-six fabrications where there were none, the check that compared a string against a pair of values and therefore could never fail, the average built from six outliers of ninety-six. Not as a confession, but because the pattern in how they were found is the most transferable thing here.
For anybody who has to decide something about an AI system and would rather decide it on evidence than on somebody's confidence.
adbl_web_anon_alc_button_suppression_t1
まだレビューはありません