『The 80,000 Hours Podcast on Artificial Intelligence』のカバーアート

The 80,000 Hours Podcast on Artificial Intelligence

The 80,000 Hours Podcast on Artificial Intelligence

著者: 80 000 Hours
無料で聴く

10 experts, 10 episodes: a crash course on transformative AI and what you can do to help shape its trajectory. This compilation features 10 key episodes of The 80,000 Hours Podcast to help listeners — particularly those new to the topic — get to grips with the potential upsides and downsides of powerful, transformative AI.80000 Hours 社会科学 科学
エピソード
  • Zero: What to expect in this series
    2026/06/05

    What might it be like to live through the creation of AI that surpasses human abilities? That future may be closer than you think.


    In this series, one expert interview at a time, we'll walk you through what's at stake — and what you could do to help.

    続きを読む 一部表示
    2 分
  • One: Will MacAskill on AI causing a “century in a decade” — and how we’re completely unprepared
    2026/06/05
    The 20th century saw unprecedented change: nuclear weapons, satellites, the rise and fall of communism, third-wave feminism, the internet, postmodernism, game theory, genetic engineering, the Big Bang theory, quantum mechanics, widespread birth control, and more. Now imagine all of it compressed into just 10 years.That’s the future Will MacAskill — philosopher, founding figure of effective altruism, and now researcher at Forethought Research — argues we need to prepare for in his paper “Preparing for the intelligence explosion.” Not in the distant future, but probably in 3–7 years.The reason: AI systems are rapidly approaching human-level capability in scientific research and intellectual tasks. Once AI exceeds human abilities in AI research itself, we’ll enter a recursive self-improvement cycle, with AI acting autonomously to create wildly more capable systems. Soon after, by improving algorithms and manufacturing chips, we’ll deploy millions, then billions, then trillions of superhuman AI scientists working 24/7 without human limitations. These systems will collaborate across disciplines, build on each discovery instantly, and conduct experiments at unprecedented scale and speed — compressing a century of progress into years.Will compares this to a mediaeval king suddenly needing to upgrade from bows and arrows to nuclear weapons to deal with an ideological threat from a kingdom he’s never heard of, while simultaneously learning he’s descended from monkeys and his god doesn’t exist.What makes this acceleration perilous is that while technology can speed up almost arbitrarily, human institutions and decision making are much more fixed.Consider the case of nuclear weapons: in this compressed timeline, there would have been just a three-month gap between the Manhattan Project’s start and the Hiroshima bombing, and the Cuban Missile Crisis would have lasted just over a day.Robert Kennedy Sr, who helped navigate the actual Cuban Missile Crisis, once said that if they’d had to make decisions faster — like in 24 hours rather than 13 days — they would likely have taken much more aggressive, much riskier actions.So there’s reason to worry about our capacity to make wise choices quickly. And in his paper, Will lays out 10 “grand challenges” we’ll need to navigate to avoid things going wrong.Will now believes we’re entering one of the most critical periods for humanity ever — with decisions made in the next few years potentially determining outcomes millions of years into the future.In this wide-ranging conversation, Will and host Rob Wiblin discuss:Why leading AI safety researchers now think there’s dramatically less time before AI is transformative than they’d previously thoughtThe three different types of intelligence explosions that occur in orderWill’s list of resulting grand challenges — including destructive technologies, space governance, concentration of power, and digital rightsHow to prevent ourselves from accidentally “locking in” mediocre futures for all eternityWays AI could radically improve human coordination and decision makingWhy we should aim for truly flourishing futures, not just avoiding extinctionLearn more and read the full transcript on the 80,000 Hours website. This episode was originally released in March 2025.Chapters:Cold open (00:00:00)Who’s Will MacAskill? (00:00:43)Why Will now just works on AGI (00:01:03)Will was wrong(ish) on AI timelines and hinge of history (00:04:21)A century of history crammed into a decade (00:09:19)Science goes super fast; our institutions don't keep up (00:16:15)Is it good or bad for intellectual progress to 10x? (00:21:44)An intelligence explosion is not just plausible but likely (00:23:41)Intellectual advances outside technology are similarly important (00:30:04)Counterarguments to intelligence explosion (00:32:42)The three types of intelligence explosion (software, technological, industrial) (00:39:00)The industrial intelligence explosion is the most certain and enduring (00:42:01)Is a 100x or 1,000x speedup more likely than 10x? (00:53:44)The grand superintelligence challenges (00:57:39)Grand challenge #1: Many new destructive technologies (01:01:29)Grand challenge #2: Seizure of power by a small group (01:09:10)Is global lock-in really plausible? (01:11:06)Grand challenge #3: Space governance (01:21:50)Is space truly defence-dominant? (01:32:19)Grand challenge #4: Morally integrating with digital beings (01:36:04)Will we ever know if digital minds are happy? (01:45:01)“My worry isn't that we won't know; it's that we won't care” (01:50:39)Can we get AGI to solve all these issues as early as possible? (01:54:05)Politicians have to learn to use AI advisors (02:07:05)Ensuring AI makes us smarter decision-makers (02:11:25)How listeners can speed up AI epistemic tools (02:15:11)AI could become great at forecasting (02:18:54)How not to lock in a bad future (02:20:26)AI takeover might happen ...
    続きを読む 一部表示
    4 時間 8 分
  • Two: Ajeya Cotra on accidentally teaching AI models to deceive us
    2026/06/05
    We don’t yet have a reliable way to tell whether an AI model is genuinely trying to help us — or faking it.A model might sincerely want to do exactly what you ask. Or it could be happy to secretly cheat, as long as its answer gets positive reinforcement during training. It might even follow the rules just to gain our trust, all while concealing goals of its own.The problem is: each of these three motivations scores the same during testing.Ajeya Cotra — previously a senior research analyst at Coefficient Giving, now working at METR (Model Evaluation & Threat Research) — explains how dangerous this dynamic could become as we train very general and very capable AI models.She likens humanity’s future trust in AI systems to an orphaned child who inherits a $1 trillion company. This child has to hire someone to run the company, guide his life, and manage his wealth — but he can only choose this person based on a work trial or interview that he designs, with no resumes or reference checks.And, because he’s so rich, all sorts of people apply — for all sorts of reasons. Some applicants will truly want to help. But the role will attract others who only pretend to care while they’re being monitored, but intend to exploit the child as soon as they can get away with it.Like a child trying to judge adults, at some point humans will need to judge the trustworthiness and reliability of machine learning models that are as goal-oriented as people, and greatly outclass us in knowledge, experience, breadth, and speed.And we can’t rely on models’ performance during training tasks to guide us, as current reinforcement learning would give the same grades to three vastly different motivations:Saints — models that genuinely care about doing what we wantSycophants — models that just want positive reinforcement for a ‘correct’ result, even if they get there with actions they know we wouldn’t approve of Schemers — models that don’t care about our interests at all, and only behave correctly as long as it serves their own agenda Worse still, training might actively encourage deception. Imagine training a model to run a business, and measuring its success by the balance in its bank account. A highly capable model might experiment with dishonest strategies. Maybe it steals some money and covers it up. (This isn’t a hypothetical worry; models often come up with creative — sometimes undesirable — approaches during training that their developers didn’t anticipate.)A model that cheats and covers its tracks would look like a star performer — and get reinforced for exactly that behaviour. If cheating is only caught some of the time, the model still might not learn to stop deceptive behaviour. Instead, it might learn that deceiving without being caught gives it a competitive advantage.In this conversation, Ajeya and host Rob Wiblin discuss the above, as well as:How to predict the motivations a neural network will develop through trainingWhether AIs in training will functionally understand that they’re AIs being trainedStories of AI misalignment that Ajeya doesn’t buyAnalogies for AI, from octopuses to aliens to can openersWhy it’s smarter to have separate ‘planning AIs’ and ‘doing AIs’The benefits of only following through on AI-generated plans that make sense to human beingsWhich approaches for fixing alignment problems Ajeya is most excited about, and which she thinks are overratedHow we might demonstrate actually scary AI failure mechanismsLearn more and read the full transcript on the 80,000 Hours website.This episode was originally released in May 2023.Chapters:Rob’s intro (00:00:00)The interview begins (00:02:38)How Ajeya’s views have changed since 2020 (00:05:09)Are neural networks more like a sped-up version of evolution, or a slower version of human learning? (00:17:42)Situational awareness (00:26:10)Misalignment stories Ajeya doesn't buy (00:42:03)The orphan heir with a trillion-dollar fortune (00:59:14)Saints, Sycophants, and Schemers (01:03:41)Ways to train safer AI systems (01:23:20)Aliens and other analogies (01:38:22)Moral patienthood (01:53:21)ARC Evaluations (01:55:35)Interpretability research (02:09:25)Rewarding models based on how good and sensible their plans seem to us (02:17:48)Overrated approaches (02:25:49)Demos of actually scary alignment failures (02:30:57)Skills to develop for doing useful work (02:37:23)Rob’s outro (02:47:24)Producer: Keiran HarrisAudio mastering: Ryan Kessler and Ben CordellTranscriptions: Katy Moore
    続きを読む 一部表示
    2 時間 50 分
adbl_web_anon_alc_button_suppression_t1
まだレビューはありません