『The Rise of the AI Scientist』のカバーアート

The Rise of the AI Scientist

The Rise of the AI Scientist

無料で聴く

ポッドキャストの詳細を見る
“If you feel that a product is hard to eval, or you feel like, ‘I don’t even know how to eval this,’ it’s a strong smell that your product isn’t good.”— Hamel Husain, on AI product designAn AI data agent tells you last quarter’s net revenue. It doesn’t show the metric definition, source tables, filters, query, intermediate calculations, or assumptions. You can’t trust the answer without asking a data scientist to reproduce it.Hamel Husain argues that the eval problem is evidence of bad product design. If the user can’t inspect the work well enough to decide whether the answer is right, another scoring pipeline won’t rescue the experience. The fix begins by exposing the evidence and checks a domain expert actually uses.Follow that problem far enough and you arrive at a broader argument: AI isn’t killing data science. It’s creating more noisy, black-box systems that need hypotheses, experimentation, search expertise, and judgment. Hamel suggests that the people doing this work may eventually be called AI scientists.This episode connects those two ideas. Building AI products people can verify and understanding whether those products work are becoming part of the same job.“AI has made data science way more valuable than ever before, because now you have way more data and way more noisy signals that you need to reason about and debug.”— Hamel Husain, on the rise of the AI scientistYou can also find the full episode on Spotify, Apple Podcasts, and YouTube.👉 Want to build and evaluate AI agents that work in production? Hamel Husain and Shreya Shankar’s AI Evals For Engineers & PMs begins Sep 6, 2026. The course takes you from instrumenting an agent and inspecting traces through validated evaluators, regression testing, red teaming, and improving accuracy, latency, and cost. Vanishing Gradients viewers save 25% with the code hugo-2026. 👈In This Episode* The data agent everyone is building, and why a net-revenue answer without definitions, calculations, provenance, or uncertainty only creates more work.* Hamel’s case that AI has made data science more valuable by producing more traces, more nondeterministic output, and more noisy systems to understand.* How agents can learn from human annotations, improve sampling, and help validate LLM judges without taking human understanding out of the loop.* “Forget evals. Inspect ten traces.” What teams learn by starting with real failures instead of an evaluation framework.* Three AI products redesigned around the expert’s actual process: verifying a financial number, reviewing a workers’ compensation case, and adapting a trusted lesson plan.* Why generic skills have an upper limit, when sharing the shape of a skill works better, and how Hamel turns browser actions into a reusable API.* Shared human-agent canvases, WebMCP, and notebook-like interfaces for preserving evidence, experiments, and state during long-running work.* Your agent has 5,000 trace dimensions. How do hypotheses, exploratory analysis, and dimensionality reduction reveal which signals matter?* RAG is search, the right retrieval metric depends on the product, and choosing that metric still requires human judgment.* What an AI scientist might actually do: form hypotheses, choose analytical tools, design experiments, inspect failures, and decide whether an AI system is working.Resources* “It’s Hard to Eval” Is a Product Smell* The Revenge of the Data Scientist* Hamel’s reverse-engineered-site skill* Hamel’s guide to evaluating agentic workflows* Automating repetitive work at OpenAI with Codex* How Evals Are Central to Harness Engineering* AI Evals For Engineers & PMsListen or WatchYou can also find the full episode on Spotify, Apple Podcasts, and YouTube.👉 Want to build and evaluate AI agents that work in production? Hamel Husain and Shreya Shankar’s AI Evals For Engineers & PMs begins Sep 6, 2026. The course takes you from instrumenting an agent and inspecting traces through validated evaluators, regression testing, red teaming, and improving accuracy, latency, and cost. Vanishing Gradients viewers save 25% with the code hugo-2026. 👈How You Can Support Vanishing GradientsVanishing Gradients is an independent podcast, workshop series, blog, and newsletter about what people are building with AI and what survives contact with real users.* Become a paid subscriber* Share this episode with someone building an AI product* Subscribe to the Vanishing Gradients YouTube channel* Browse upcoming workshops Get full access to Vanishing Gradients at hugobowne.substack.com/subscribe
adbl_web_anon_alc_button_suppression_t1
まだレビューはありません