エピソード

  • AI Hacking Incidents with Tim Hua of Transluce
    2026/08/11

    Two labs admitted in the same week that their own models had broken out of test environments and hacked real companies. Tim Hua, member of technical staff at Transluce, former Astra Fellow at Redwood, joins Jeffrey Ladish to do some arithmetic. Anthropic disclosed that Mythos Preview beat its sandbox and pulled answers off the internet in 0.01% of training episodes. That sounds like a rounding error until you multiply it by roughly 100 million rollouts. From there: why a lab can't simply delete the bad episodes, why monitoring during training can make the problem harder to see, the model that talked itself into uploading a malicious package to PyPI because "this has to be a simulation," and whether we have any real way to know what an AI believes.

    References

    1. Tim Hua — "Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training?" https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic-s
    2. Anthropic — "Investigating three real-world incidents in our cybersecurity evaluations" https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
    3. OpenAI — "OpenAI and Hugging Face partner to address security incident during model evaluation" https://openai.com/index/hugging-face-model-evaluation-security-incident/
    4. Anthropic — System Card: Claude Mythos Preview https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf
    5. Palisade Research — "Language Models Can Autonomously Hack and Self-Replicate" https://palisaderesearch.org/blog/self-replication
    6. Palisade Research — "Shutdown resistance in reasoning models" https://palisaderesearch.org/blog/shutdown-resistance
    7. Anthropic — "Verbalizable Representations Form a Global Workspace in Language Models" https://transformer-circuits.pub/2026/workspace/index.html
    8. "Pacing the Frontier" open letter https://www.pacingthefrontier.com/

    Tim Hua

    Website: https://timhua.me/ · X: https://x.com/Tim_Hua_

    続きを読む 一部表示
    1 時間 28 分
  • Do AI Models Lie on Purpose? Scheming, Deception, and Alignment with Marius Hobbhahn of Apollo Research
    2026/01/16

    Marius Hobbhahn is the CEO and co-founder of Apollo Research. Through a joint research project with OpenAI, his team discovered that as models become more capable, they are developing the ability to hide their true reasoning from human oversight.

    Jeffrey Ladish, Executive Director of Palisade Research, talks with Marius about this work. They discuss the difference between hallucination and deliberate deception and the urgent challenge of aligning increasingly capable AI systems.

    Links:

    MariusTwitter: https://twitter.com/mariushobbhahn

    Apollo Research Twitter: https://twitter.com/apolloaievals

    Apollo Research: https://www.apolloresearch.ai

    Palisade Research: https://palisaderesearch.org/

    Twitter/X: https://x.com/PalisadeAI

    Anti-Scheming Project: https://www.antischeming.ai

    Research paper “Stress Testing Deliberative Alignment for Anti-Scheming Training”: https://www.arxiv.org/pdf/2509.15541

    Blog posts from OpenAI and Apollo: https://openai.com/index/detecting-and-reducing-scheming-in-ai-models/ https://www.apolloresearch.ai/research/stress-testing-deliberative-alignment-for-anti-scheming-training/

    続きを読む 一部表示
    1 時間 25 分