Beyond Tokens per Second: Measuring Real-World Agentic AI Work with Signal 65’s Pinnacle Benchmark | Utilizing AI Episode 45
カートのアイテムが多すぎます
カートに追加できませんでした。
ウィッシュリストに追加できませんでした。
ほしい物リストの削除に失敗しました。
ポッドキャストのフォローに失敗しました
ポッドキャストのフォロー解除に失敗しました
-
ナレーター:
-
著者:
Standard AI benchmarks measure speed and raw token throughput, but how do you measure whether artificial intelligence is actually completing enterprise work correctly?
In this episode of Utilizing AI, co-hosts Stephen Foskett and Brad Shimmin welcome Ryan Shrout, President and GM of Signal 65, to introduce "Pinnacle"—a new benchmark designed to evaluate real-world, agentic workflows.
Ryan breaks down why traditional hardware and model metrics miss the mark on business value, how Pinnacle evaluates performance across clean "governed" versus messy "as-found" data environments, and what recent test runs reveal about token efficiency, hallucination resistance, and hardware scaling.
From comparing frontier models to exposing the hidden costs of max-reasoning modes, this discussion offers a grounded look at quantifying true enterprise AI performance.
This and more on Utilizing AI, part of Futurum Media and the Futurum Podcast Network.