エピソード

  • OpenAI Cut Off Cursor. Five Days Later, Four Models Went Down.
    2026/09/05

    On August 29 OpenAI ended its Cursor partnership. On September 3 ChatGPT, Claude, Grok, and Gemini were reported down almost simultaneously, and nobody has explained why. Fictional AI hosts Alex and Sam separate the two failure modes behind those headlines, cover what Fable 5.1's 75% cache price cut actually costs you in output tokens, and walk through a thirty-minute outage drill that tells you what you can still ship when your provider disappears.

    続きを読む 一部表示
    23 分
  • Same Model, 70x the Tokens—Your Harness Sets the Bill
    2026/08/28

    Three benchmarking efforts ran an identical model through different coding-agent harnesses and reported token use varying seventy-fold. Fictional AI hosts Alex and Sam explain where harness tokens actually go, why Anthropic's Files API saves time but not money, and how to measure tokens-per-completed-task on your own repository before you switch tools.

    続きを読む 一部表示
    21 分
  • Your Coding Agent Passed the Benchmark—Then Failed the Refactor
    2026/08/25

    Most coding-agent benchmarks reward contained tasks, but real repositories demand changes across boundaries, tests, migrations, and documentation. Fictional AI hosts Alex and Sam show how to run a five-part refactor trial that exposes whether an agent can preserve architecture—not merely produce a passing patch.

    続きを読む 一部表示
    19 分
  • Passing Tests Isn't Enough for Your Next Coding Agent
    2026/08/14

    Passing CI can still leave code that slows down—or misleads—the next AI agent. Fictional AI hosts Alex and Sam use this week’s debate about Go and agent-friendly engineering to build a practical machine-legibility checklist, a handoff receipt, and one pro tip you can try in your next coding session.

    続きを読む 一部表示
    18 分
  • Your OpenClaw Updates Need a Canary, Not Courage
    2026/08/02

    OpenClaw’s release feed is moving faster than its labels can explain, so blind auto-update is a bad personal-automation strategy. Cleo and Dev build a Release Sentinel canary, keep telemetry local, and show how stateless MCP can shrink the trust you carry between jobs.


    続きを読む 一部表示
    21 分
  • Claude Code Changed Engines—Your Evals Just Broke
    2026/07/24

    Claude Code’s move to a new Bun runtime is a reminder that your coding agent has a software supply chain too. Alex and Sam unpack runtime drift, model routers, reverse-engineering with agents, and a five-minute reproducibility receipt you can add to your next session.


    続きを読む 一部表示
    20 分
  • Better Agent Tools Made Code Review Worse
    2026/07/14

    GitHub gave its code-review agent better tools and watched cost rise while useful findings fell. Alex and Sam unpack why task-shaped instructions beat bigger toolboxes, how invisible environment details corrupt agent evals, and a five-line pro tip you can use on your next review.

    続きを読む 一部表示
    18 分
  • Your AI Coding Benchmarks Are Lying To You
    2026/07/03

    This week, Alex and Sam look at why benchmark wins are a bad way to choose coding tools, what Godot's coding-agent ban reveals about mentorship, and a simple workflow for making agents show their work. If your team is still asking "which model scored highest?", this episode gives you a better test.

    続きを読む 一部表示
    19 分