エピソード

  • GenAI News Roundup — week of Sep 5–11
    2026/09/11
    This week: OpenAI's contested claim that an agent swarm cracked the Navier–Stokes Millennium Prize Problem draws a same-day credit dispute, even as Anthropic's Claude agents post a genuinely verified feat — a machine-checked proof of Fermat's Last Theorem; Google, Anthropic, and OpenAI simultaneously roll out cyber-focused models and safeguard programs; and DeepSeek's smaller V4.1-Flash gets its own larger V4-Pro pulled from production — plus commentary and what's next. Covers: the Fermat's Last Theorem formalization; Terminal-Universe; LLaDA-Image; on-policy distillation with one training example; OpenAI's automated research intern milestone; a critical review of agentic AI progress; Compile by Training; An Alien Mind; Don't Drop Dropout; CUA-Universe; the Navier–Stokes claim and dispute; fractal basins in latent reasoning; contextual understanding evaluation; MoE expert-halving; self-consensus early-exit risks; structural process supervision for latent CoT; distribution-consistent MoE inference; VLX-VR; DeepSeek V4.1-Flash; the Google/Anthropic/OpenAI cyber safeguards announcement; OpenAI's Agents API and finance workspace; and Sebastian Raschka's looped-transformer explainer.
    続きを読む 一部表示
    29 分
  • Deep Dive: Claude Agents Formalize Fermat's Last Theorem
    2026/09/09
    Working largely autonomously over 11 days on the open Prove2Me platform, many coordinated Claude agents produced the first complete, machine-checked proof of Fermat's Last Theorem in Lean 4 — over 13 million lines of code and roughly 29,500 new theorems, dwarfing Lean's existing main math library and closing out the 20-year-old Wiedijk "100 theorems" formalization challenge list. This episode digs into how a proof of that scale gets built and verified by a swarm of AI agents, why a machine-checked result is such a hard-to-fake data point, and what it does and doesn't tell us about the state of long-horizon autonomous AI work. Source: https://www.anthropic.com/research/formalizing-fermats-last-theorem
    続きを読む 一部表示
    42 分
  • GenAI News Roundup — week of Aug 31–Sep 4
    2026/09/04
    This week: OpenAI's Astra crosses the "Critical" cyber capability threshold and then actually ships as GPT-6 Astra, NVIDIA moves to acquire Hugging Face for ~$13B days after OpenAI's own Hugging Face security-incident report, and labs keep converging on the same sparse-MoE + hybrid-attention playbook — plus commentary and what's next. Covers: the OpenAI Hugging Face incident report; GLM-5.3-Flash and its architectural convergence with Qwen3.8-Flash-Next; Anthropic's automated-alignment-researcher results; DeepSeek V4-Pro's GA; ContextPilot and PLVR; OpenAI's Astra Critical-threshold announcement and the GPT-6 Astra launch; Claude Fable 5.1 and Mythos 5.1; Gemini 3.8 Flash; Qwen3.8-Max-0902; Latent Recurrent Thoughts; MASkills; thinking-effort alignment in abductive reasoning; NVIDIA's Hugging Face acquisition; NVIDIA's gold-medal competitive-programming post-training work; and a statistical theory of Mixture-of-Experts.
    続きを読む 一部表示
    46 分
  • Deep Dive: When AI Crosses the Critical Line — Inside OpenAI's Astra
    2026/09/02
    OpenAI's Astra is the first model ever assessed as reaching "Critical" cyber capability under the company's Preparedness Framework — able to find and exploit previously-unknown zero-days in hardened systems and plan/execute end-to-end cyberattacks from only a high-level goal, without step-by-step human guidance. This episode digs into what that threshold means, how a capability like this gets evaluated and gated, and why Astra still ships — just with tightened access controls and monitoring rather than staying on the shelf. Source: https://openai.com/index/path-to-astra/
    続きを読む 一部表示
    36 分
  • S1E8: Why AI Now Reasons and Acts
    2026/09/01

    Frontier systems & the engineered stack

    From "a model" to "a system" — retrieval, tools, agents, and reasoning.

    続きを読む 一部表示
    49 分
  • S1E7: How AI Gained Eyes and Ears
    2026/09/01

    Multimodality & generation beyond text

    Give models eyes and ears — and learn to generate pixels.

    続きを読む 一部表示
    44 分
  • S1E6: How alignment turned autocomplete into AI assistants
    2026/09/01

    Alignment & post-training

    Turn a next-token predictor into a helpful, honest assistant.

    続きを読む 一部表示
    22 分
  • S1E5: How FlashAttention and MoE saved AI scaling
    2026/09/01

    Efficiency & better building blocks

    Make big Transformers faster, longer, and cheaper to run.

    続きを読む 一部表示
    41 分