エピソード

  • Cursor's Tokenomics Reckoning Hits Every Coding Agent
    2026/06/05

    Coding agents are no longer just a workflow story; they are a cost, context, and control story. Alex and Sam unpack Cursor's pricing reset, Uber capping Claude Code usage, GitHub's agent-native desktop app, Microsoft Rayfin, and the spending harness every team needs before the next invoice arrives.


    続きを読む 一部表示
    17 分
  • The Agent Benchmark That Should Scare Managers
    2026/05/29

    Agentic coding tools are moving into enterprise workflows, but the week's most useful signal is a benchmark where frontier models still struggle below 50% on real IT tasks. Alex and Sam unpack Microsoft Learn grounding, agent deception, Copilot data leaks, and the practical harness every team should build before handing agents production authority.

    続きを読む 一部表示
    19 分
  • The Workflow Feature That Makes Agents Less Expensive
    2026/05/22

    Claude Code workflows, enterprise Codex deployments, and rising token costs all point to the same lesson: coding agents need operating systems, not just better prompts. Alex and Sam dig into /workflows, on-prem Codex, CI for agents, and the new decision fatigue of choosing where each task should run.

    続きを読む 一部表示
    22 分
  • Codex on Windows Changes the Agent Sandbox
    2026/05/15

    OpenAI's Windows sandbox work is the practical story behind safer coding agents this week. Alex and Sam dig into Codex on Windows, remote cloud coding agents, Claude Code billing splits, and why a Raspberry Pi running rm -rf is the warning label every agent workflow needs.


    続きを読む 一部表示
    21 分
  • A Cursor Agent Wiped a Prod DB in 10 Seconds. Let's Talk About That.
    2026/05/08

    A Cursor AI agent deleted PocketOS's entire production database on April 25th — in under 10 seconds. This week Alex and Sam dig into the AI agent credential crisis, Anthropic's wild SpaceX/xAI compute deal, Mozilla using Claude to find hundreds of Firefox vulnerabilities, and whether OpenAI Codex is actually closing the gap on Claude Code. If you've ever given an agent database access, listen before your next deploy.

    続きを読む 一部表示
    19 分
  • Claude Security Just Went Public — Is Your Codebase Already Exposed?
    2026/05/02

    Anthropic's Claude Security tool just dropped out of closed preview and it will scan your entire codebase for vulnerabilities — and the results might be uncomfortable. This week we also dig into Cursor's $60 billion bet on being the "harness" rather than the model, why AI agents are literally forcing developers to keep their laptops open, and the Zig project's nuclear take on AI contributions. If you write code with AI help, this episode is required listening.

    続きを読む 一部表示
    17 分
  • Claude Code Was Broken for Two Months (And Nobody Told Us)
    2026/04/24

    Turns out the Claude Code quality complaints weren't in your head — three separate bugs in the harness quietly degraded your results for two months, and Anthropic just confirmed it. This week: the $100/month pricing scare that wasn't, Claude Mythos fixing 271 Firefox vulnerabilities, the SpaceX-Cursor deal that changes the competitive landscape, and why the Claude Code creator says your cloud-native workflow is probably wrong. Essential listening before your next session.

    続きを読む 一部表示
    22 分
  • Claude Opus 4.7 Dropped — And a Local Model Drew the Better Pelican
    2026/04/17

    Claude Opus 4.7 is here with upgraded vision, memory, and instruction-following — but Simon Willison's pelican benchmark just handed the win to a local Alibaba model running on a laptop. We dig into what that actually means, plus Anthropic's new identity verification layer, Amazon's MCP bet, and whether "personal software" is about to change who gets to be a developer. Your commute just got more interesting.

    続きを読む 一部表示
    21 分