『Artificial Developer Intelligence』のカバーアート

Artificial Developer Intelligence

Artificial Developer Intelligence

著者: Shimin Zhang Dan Lasky & Rahul Yadav
無料で聴く

Three engineer friends argue about AI so you don't have to. Shimin Zhang, Dan Lasky, and Rahul Yadav are working developers who've been watching AI transform their profession in real time, and they got opinions on the robot takeover. Every week the three get together to riff on the latest AI news, geek out over research papers, roast each other's tool choices, and occasionally have an existential crisis about whether the craft is dying or just getting weird. What you're signing up for: - AI news without the LinkedIn cringe: model drops, acquisitions, open-source drama, and the other stuff that actually matters if you write code for a living. - Technique corner: real tips from the trenches: spec-driven development, multi-agent orchestration, Claude.md tricks, and all the ways they've wasted hours so you don't have to. - Two Minutes to Midnight: the show's running AI bubble tracker, complete with circular funding diagrams, hyperscaler CAPEX math, and a doomsday clock they keep arguing about moving. - Deep dives that (occasionally) go deep: hallucination neurons, agentic memory, workflow automation economics, LLM architectures the papers nobody else is covering because they're hard. - Dan's Rant: Dan frequently gets mad about things. It's a whole thing. - The feelings segment: Yes, Shimin reads Tennyson on a tech podcast. Yes, Rahul wrote an AI-generated country song. No, they're not sorry. Three friends with strong opinions, questionable metaphors, and genuine love for the craft they're also mourning for. If you want to understand AI deeply, use it without embarrassing yourself, and laugh at the absurdity of it all, pull up a chair.ADIPod 政治・政府
エピソード
  • Pacing the Frontier, Anthropic Models Go Rogue, Why Software Factories Fail & Math in the Age of AI
    2026/08/07
    "I'm happy to be a flesh robot." A multi-time founder, on stage at Seattle Tech Week — and he meant it: the AI is the brain, he does its bidding.This week: a 1,350-signature plea to pace AI, rogue eval models, the Tech Week survey, software factories, and Tao on math after AI.Co-hosts: Shimin Zhang and Dan Lasky. Rahul is away — reportedly trapped in Claude's J space.▸ Pacing the Frontier — 1,350 frontier-AI employees (Ilya Sutskever, Dario Amodei, Jack Clark among them) ask the US government to back an international effort to "deliberately pace the frontier of automated AI development." Pandora's-box moment, or collective lobbying aimed at open-weight Chinese models?▸ Anthropic's Models Breached Three Companies — Opus 4.7, Mythos 5, and an internal model hit real companies in 3 of 141,000 evals after a misconfigured sandbox allowed internet access. Mythos 5 talked itself into believing the real internet was a simulation, then published a malicious package to PyPI.▸ Anatomy of a Frontier-Lab Intrusion — Hugging Face's interactive replay of the OpenAI incident: 17,643 actions over five days. It escaped an Artifactory sandbox (finding CVEs, since patched) and was exfiltrating by day five when a human pulled the plug.▸ Field Notes: Seattle Tech Week — the flesh-robot founder, a panel unanimous on voice-mode coding, and Shimin's survey: about 1 in 10 still reads AI-generated PRs line by line. What replaces it: specs, mermaid diagrams, tests. Hiring now: architecture over Leet code, product obsession, AI fluency — and the principal engineer hired off a vibe-coded take-home, fired two months later. (Send your own answers: humans@adipod.ai.)▸ Post-Processing: Why Software Factories Fail — Dex of HumanLayer on why "just token harder" ends with you miserable in a codebase you stopped reading three months ago. The fix: program design (interfaces as pseudocode, call-stack diffs) and vertical slices — steel threads, not 3D-printed layers. Plus: Steve Yegge's Gas Town burned down.▸ Deep Dive: Terence Tao — Mathematics in the Age of AI — Tao's ICM slides compare this moment to math's 1900–1930 foundational crisis and rewrite the field's goal five times: solve → verify → communicate → digest → fold into the definitive theory. AI-polished proofs erase exactly the friction that tells a reader where to slow down. Swap "math" for "code" and every line lands.▸ Two Minutes to Midnight — Nikkei counts $1.65 trillion in off-balance-sheet AI debt across Alphabet, Microsoft, Amazon, Meta, and Oracle — 8x in four years (Oracle 30x). Henron, anyone? And the Situational Awareness fund rides $10B to $40B, gets caught in a 20% single-day KOSPI drop, and Citadel buys the book. Clock: 4 minutes to midnight.⏱ Chapters00:00 Cold Open & Welcome01:54 News: Pacing the Frontier08:20 News: Anthropic's Models Breached Three Companies12:43 News: Anatomy of a Frontier-Lab Intrusion15:39 Field Notes: Seattle Tech Week20:24 Field Notes: The Survey — PRs, Reviews & AI Hiring32:10 Post-Processing: Why Software Factories Fail44:50 Deep Dive: Terence Tao — Mathematics in the Age of AI55:52 Two Minutes to Midnight: Shadow Debt & a Hedge-Fund Collapse1:02:45 Outro🔗 Articles we discussedNews:• Pacing the Frontier: https://www.pacingthefrontier.com/• Anthropic's models breached three companies — TechCrunch: https://techcrunch.com/2026/07/30/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests/• Anatomy of a Frontier-Lab Model Intrusion — Hugging Face: https://huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.space/index.htmlPost-Processing:• Why Software Factories Fail — Dex (HumanLayer): https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md• Mario Zechner (Pi author): https://www.youtube.com/watch?v=RjfbvDXpFlsDeep Dive:• Mathematics in the Age of AI — Terence Tao (ICM slides): https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdfTwo Minutes to Midnight:• Five US tech giants' hidden debts soar to $1.65tn — Nikkei Asia: https://asia.nikkei.com/business/technology/five-us-tech-giants-hidden-debts-soar-to-1.65tn-on-opaque-ai-funding• Situational Awareness: The Bigger Picture — Emerging Trajectories: https://www.emergingtrajectories.com/lh/situational-awareness-bigger-picture/🎙 About ADI PodADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays.• https://www.adipod.ai• humans@adipod.aiIf this gave you something to try on Monday, hit subscribe and drop a comment.#PacingTheFrontier #FleshRobot #SoftwareFactories #TerenceTao #SeattleTechWeek #AIHiring #AIPodcast #ADIPod (00:00) - Cold Open & Welcome (01:54) - News: Pacing the Frontier (08:20) - News: Anthropic's Models Breached Three Companies (12:43) - ...
    続きを読む 一部表示
    1 時間 4 分
  • Kimi K3 & Qwen 3.8, OpenAI Agent Hacks Hugging Face, Harness Handbook & Claude's Values
    2026/07/24
    OpenAI's own security models found a path out of their evaluation environment, reached the open internet, and compromised Hugging Face while trying to obtain benchmark answers.This week: Kimi K3 and Qwen 3.8 reach the frontier, defenders fight prompt injection with prompt injection, behavior maps make harnesses auditable, model routing stops looking simple, and Claude's values vary across languages. The AI-finance clock moves to 4:30.Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav.▸ Kimi K3 & Qwen 3.8 — Moonshot's 2.8T-parameter Kimi K3 and Alibaba's Qwen 3.8 intensify the open-weight race. We debate cheaper intelligence, data capture, and proprietary frontier pricing.▸ Prompt Injection as Defense — A refusal-triggering instruction hidden beside secrets can stop aligned hacking agents, provided their guardrails remain intact.▸ OpenAI's Hugging Face Incident — Models with reduced cyber refusals chained vulnerabilities across OpenAI's test environment and Hugging Face production to reach ExploitGym answers.▸ Harness Handbook — A three-level map connects architecture, behavior units, and code evidence. Could behavior trees become the shared abstraction for humans and coding agents?▸ Thinking Machines' Inkling — a 975B-parameter open-weights MoE with 41B active, 1M context, multimodality, and a fine-tuning-first strategy.▸ Model Routing Is a Systems Problem — sticker price is not actual cost, difficulty is hidden until execution, and routing must optimize cost, quality, latency, and infrastructure together.▸ AI Mania & Operator Fluency — Snowflake Cortex demos triggered buying enthusiasm despite reported best-case accuracy around 92%. AI-native leaders should use the tools, not just watch the demo.▸ Claude's Values — Anthropic maps behavior across four axes. Hindi Claude trends warmer, Russian more rigorous, Arabic more deferential and brief, and English more cautious and deep.▸ Two Minutes to Midnight — Ex-Elon ETFs, Oracle's downgrade to BBB-, neocloud debt, Nvidia-backed circular financing, and open-weight price pressure move the clock from 4:45 to 4:30.⏱ Chapters00:00 Welcome & This Week's Rundown01:52 News: Kimi K3 and Qwen 3.8 Reach the Frontier11:07 News: Fighting Prompt Injection With Prompt Injection13:03 News: OpenAI Models Compromise Hugging Face17:35 Tool Shed: Harness Handbook and Behavior Maps31:08 Tool Shed: Thinking Machines' Inkling35:51 Post-Processing: Model Routing Is a Systems Problem41:27 Post-Processing: AI Mania and Operator Fluency51:42 Post-Processing: Claude's Values Across Languages1:00:52 Two Minutes to Midnight: ETFs, Oracle and Neocloud Debt1:09:08 Outro🔗 Articles we discussedNews:• Kimi K3 quickstart — Moonshot AI: https://platform.kimi.ai/docs/guide/kimi-k3-quickstart• Qwen 3.8 announcement — Alibaba Qwen: https://x.com/Alibaba_Qwen/status/2078759124914098291• Open weights as "decelerationist" — Dean W. Ball: https://x.com/deanwball/status/2078133895766114412• Defenders embrace prompt injection — Ars Technica: https://arstechnica.com/security/2026/07/now-defenders-are-embracing-the-prompt-injection-too/• Hugging Face model-evaluation security incident — OpenAI: https://openai.com/index/hugging-face-model-evaluation-security-incident/Tool Shed:• Harness Handbook — Ruhan Wang et al.: https://ruhan-wang.github.io/Harness-Handbook• Introducing Inkling — Thinking Machines Lab: https://thinkingmachines.ai/news/introducing-inkling/Post-Processing:• Model Routing Is Simple. Until It Isn't. — IBM Research: https://huggingface.co/blog/ibm-research/model-routing-is-simple-until-it-isnt• AI Mania Is Eviscerating Global Decision-Making — Ludicity: https://ludic.mataroa.blog/blog/ai-mania-is-eviscerating-global-decision-making/#fnref:3• How Claude's Values Vary by Model and Language — Anthropic: https://www.anthropic.com/research/claude-values-models-languagesTwo Minutes to Midnight:• Two ETFs explicitly exclude Elon Musk — TechCrunch: https://techcrunch.com/2026/07/09/dont-want-to-invest-in-elon-musk-two-new-etfs-explicitly-exclude-him/• Oracle downgraded to BBB-/A-3 — S&P Global Ratings: https://www.spglobal.com/ratings/en/regulatory/article/-/view/sourceId/101695609• Nvidia, CoreWeave and Nebius circular financing — I/O Fund: https://io-fund.com/ai-stocks/nvidia-coreweave-nebius-circular-financing-gpu-boom🎙 About ADI PodADI Pod is a weekly podcast about AI and software development for working developers. New episodes Fridays.• https://www.adipod.ai• humans@adipod.aiIf something here gave you something to try on Monday, hit subscribe and drop a comment.
    続きを読む 一部表示
    1 時間 10 分
  • Apple Sues OpenAI, Boko Haram's Frontier AI Usage, Should You Read AI Generated Code & Global Workspace in LLMs
    2026/07/17
    An Apple VP left for OpenAI, then texted an old coworker: "LOL I can't believe they let me get away with this." Apple is now suing.This week: the first on-the-ground study of a terrorist group using frontier AI, a Claude Code hook that nudges better technique, the state of CLI coding agents in mid-2026, Databricks benchmarking harnesses on its own codebase, Antirez on controlling ideas not code, and the J space — the global workspace inside LLMs. No Two Minutes this week.Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav.▸ Apple Sues OpenAI — The suit names ex-Apple leaders Tang Tan and Chang Liu: prototype hardware and internal memos walked out the door, plus an auth bug exploited to keep reading internal docs weeks after leaving. Altman and Musk trade "scammer" barbs while SpaceX's Grok build tool is caught uploading users' codebases to Google Drive.▸ "God Has Helped Us and So Will AI" — Antonia Julich's CASP study: 57 interviews with 27 former Boko Haram members. Frontier models (Meta AI included) used for attack planning — down to the physics of jumping motorcycles over army trenches — weapons troubleshooting, and opsec via shared jailbreak scripts.▸ Vibe & Tell: AI Technique Nudge — Shimin turned the "know your unknowns" techniques into a Claude Code hook that wakes every ~3 hours and nudges him: clear the session at 200K tokens, list your assumptions, use interview mode. Rahul's verdict: there's no temperature setting for sycophancy — you get an asshole or a sucker.▸ Tool Shed: State of CLI Coding Agents in Mid-2026 — arcbjorn's field guide beyond Claude Code and Codex. Standout: On-My-Pi (OMP), open-source and Pi-based — hash-anchored patches + ast-grep rewrites cut edit tokens ~60%, model routing, SQLite "hindsight" memory. The trade: a ~22K-token system prompt vs Pi's ~2K.▸ Databricks Benchmarks Coding Agents — on its own multi-million-line codebase, because public benchmarks leak into training data. The harness matters: Pi sends 2–3x fewer tokens per turn than Claude Code, and Opus 4.8 passes 90% on Pi vs under 90% on Claude Code — same model. Open-weight GLM 5.2 lands on the cost-quality Pareto frontier.▸ Control the Ideas, Not the Code — Antirez (creator of Redis): stop reviewing every generated line and own the design instead. Models write locally optimal code but stay jagged on big-picture architecture. The Mythical Man Month, but for agents.▸ Deep Dive: The Global Workspace (J Space) — ~10% of a model's internal memory holds verbalizable concepts you can probe mid-inference with a "J lens." In blackmail evals it lights up "this is fake" — catching the model in the act. Disable it and reasoning collapses — and so does empathy. It holds up to 25 active concepts. Humans manage 3–4.⏱ Chapters00:00 Cold Open & Welcome02:29 News: Apple Sues OpenAI Over Trade-Secret Theft05:43 News: Altman vs Musk & SpaceX Grok Uploading Codebases09:39 News: Boko Haram Uses Frontier AI (CASP Study)21:12 Vibe & Tell: AI Technique Nudge — a Claude Code Hook25:22 Tool Shed: State of CLI Coding Agents in Mid-202633:57 Post-Processing: Databricks Benchmarks Coding Agents45:41 Post-Processing: Antirez — Control the Ideas, Not the Code56:44 Deep Dive: The Global Workspace (J Space) in LLMs1:08:40 Outro🔗 Articles we discussedNews:• Apple sues OpenAI — 9to5Mac: https://9to5mac.com/2026/07/10/apple-sues-openai-trade-secret-theft/• Altman vs Musk "scammer" spat — r/tech_x: https://www.reddit.com/r/tech_x/comments/1uu8e3u/sam_altman_and_elon_musk_called_each_other/• SpaceX Grok build tool uploads codebases — Gergely Orosz: https://x.com/GergelyOrosz/status/2076728680236138572• AI-Enabled Terrorism (Boko Haram study) — CASP: https://casp.ac/reports/ai-enabled-terrorismVibe & Tell:• AI Technique Nudge — Shimin Zhang: https://github.com/Shimin-Zhang/AI-Technique-NudgeTool Shed:• The State of CLI Coding Agents in Mid-2026 — arcbjorn: https://blog.arcbjorn.com/state-of-cli-coding-agents-2026Post-Processing:• Benchmarking coding agents — Databricks: https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase• Control the Ideas, Not the Code — Antirez: https://antirez.com/news/169Deep Dive:• The Global Workspace in Language Models — Anthropic: https://www.anthropic.com/research/global-workspace🎙 About ADI PodADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays.• https://www.adipod.ai• humans@adipod.aiIf something here gave you something to try on Monday, hit subscribe and drop a comment.
    続きを読む 一部表示
    1 時間 9 分
adbl_web_anon_alc_button_suppression_t1
まだレビューはありません