エピソード

  • The Psychology of Software Teams with Dr. Cat Hicks
    2026/09/11
    Engineering teams that lowered production pressure to make space for learning ended up roughly 18 months ahead of everyone else.This week is different: no news, no clock. Shimin and Dan sit down with Dr. Cat Hicks — psychologist for the humans in tech, principal scientist at Catharsis, and author of The Psychology of Software Teams — on how organizations misunderstand developer cognition, how to learn deliberately with AI instead of letting your skills atrophy, and what a healthy AI-native team actually looks like.Hosts: Shimin Zhang and Dan Lasky. Guest: Dr. Cat Hicks.▸ Brains in Jars — Cat's name for organizations that treat developers as fungible, disembodied cognition. Developers are highly valued — just stripped of their humanity. It isn't even an accurate model: real problem-solving is heavily social. Brains in jars never produced great software.▸ The 10x Engineer Myth — The lone-genius stereotype sticks because it rewards grind-set self-belief; Cat argues its costs are far bigger than we like to see. Dan brings the McCarthy core-protocols check-in ritual (mad, sad, glad, or afraid), and Cat explains why a technically credible engineer showing feelings rewires a team's stereotypes.▸ Effortful Learning & the learning-opportunities Skill — Cat's Claude/Codex skill generates 10-15 minute deliberate-practice exercises from the files you're already working in. Built against the deficit mindset ("our brains will melt"). Ten minutes does something to your mind. It doesn't have to be five hours.▸ The Fluency Illusion & Metacognitive Decoupling — We are bad judges of our own learning: fluent-feeling output masks understanding that isn't there. Cat on the scary lab studies: do you ever remember what you copy-paste? What works: self-quizzing, sketching architecture before implementing, interrupting agentic sessions with a short learning op. Metacognition is one of the few tractable interventions — working memory is hard to change, metacognitive strategy isn't, and it predicts life success.▸ Measuring Your Own Skills — No good instruments exist for working developers to track their own skills over time. Cat, a former assessment scientist, wants to build them — which might accidentally fix interviewing too.▸ Healthy AI-Native Teams — Dumping unreviewed AI output on teammates is a cultural statement, not a technical decision: it says their problem-solving doesn't matter. The fix is a shared team mental model — decide explicitly how you'll use AI for the next month, then test whether it worked. Teams that lowered overproduction pressure ended up roughly 18 months ahead. Scale in production demands matching scale in testing and auditing — borrow probabilistic thinking from computational biology and robotics.⏱ Chapters00:00 Cold Open — No News, No Clock00:35 Meet Dr. Cat Hicks01:14 Is Software Engineering Abnormal Psychology?03:27 Brains in Jars — Why Orgs Misread How Developers Think06:43 The 10x Engineer Myth09:30 The McCarthy Core-Protocols Check-In12:45 Effortful Learning & the learning-opportunities Skill18:50 The Fluency Illusion & Metacognitive Decoupling22:51 Measuring Your Own Skills (and Fixing Interviews)24:38 What Are We Actually Learning Right Now?26:25 Informed Patient — the Same Problem Outside Tech29:39 What a Healthy AI-Native Team Looks Like33:54 The Book, the Podcast & the Newsletter🔗 LinksThe book:• The Psychology of Software Teams — Routledge: https://www.routledge.com/The-Psychology-of-Software-Teams/Hicks/p/book/9781032963389Cat's work:• Dr. Cat Hicks: https://www.drcathicks.com/• Catharsis: https://www.catharsisinsight.com/• Change, Technically (podcast): https://www.changetechnically.fyi/• Fight for the Human (newsletter): https://www.fightforthehuman.com/The skills:• learning-opportunities: https://github.com/DrCatHicks/learning-opportunities• learning-goal: https://github.com/DrCatHicks/learning-goalFollow Cat:• Bluesky: https://bsky.app/profile/grimalkina.bsky.social• LinkedIn: https://www.linkedin.com/in/drcathicks/🎓 About Dr. Cat HicksDr. Cat Hicks is a psychologist for the humans in tech and principal scientist at Catharsis. She wrote The Psychology of Software Teams (CRC Press, July 2026), co-hosts Change, Technically, writes the Fight for the Human newsletter, created the learning-opportunities and learning-goal skills, and authored the Developer Thriving and AI Skill Threat frameworks.🎙 About ADI PodADI Pod (Artificial Developer Intelligence) is a weekly conversation show for working developers, sorting the hype around AI from what it actually delivers. New episodes Fridays.• https://www.adipod.ai• humans@adipod.aiIf the fluency illusion hit a little too close to home, subscribe and tell us what you're doing about it.#CatHicks #PsychologyOfSoftwareTeams #Metacognition #10xEngineer #AICoding #AIPodcast #ADIPod
    続きを読む 一部表示
    36 分
  • Meta's AI Backfires, GLM 5.3 Flash, OpenAI's Jalapeño Chip & the End of Programming
    2026/09/04
    One developer used Fable V to rewrite Bun in Rust: 1M+ lines, ~7,000 commits, 11 days, ~$165K in API costs. It now underpins Claude Code.This week: Meta's AI-native plan ends in 40% more incidents and a cancelled layoff round, OpenAI cuts Cursor off after the SpaceX acquisition, our new Model Review segment debuts with ZAI's GLM 5.3 Flash, OpenAI's Jalapeño chip posts absurd tokens-per-kilowatt numbers, and the AI bubble clock moves backward for the first time in a month.Hosts: Shimin Zhang, Dan Lasky, and Rahul Yadav — back from the Galápagos, where everyone talks in geological time.▸ News: Meta's AI-Native Plan Backfires — Project OT ("organizational transformation") was meant to cut headcount 25% by replacing workers with AI agents. Code changes rose 220% year over year; features reaching users rose only 36%; major technical and security incidents rose 40%, with firefighting time up as much as 70% — and the second round of layoffs quietly died. The supercar analogy applies: the tech wasn't the bottleneck, the organization was.▸ News: OpenAI Cuts Cursor Off — After SpaceX's ~$60B Cursor acquisition, OpenAI will pull its models by November 12, citing terms-of-service history — and, without quite saying it, distillation. Consolidation begets counter-moves; we ask what Cursor's moat is that OpenRouter doesn't already have.▸ Model Review (new segment): GLM 5.3 Flash — OpenRouter's free mystery model "Ox Alpha" turns out to be ZAI's GLM 5.3 Flash: above Sonnet 5 at max thinking and Grok 4 on the Artificial Analysis index, at $0.075 per million input tokens, served entirely on Chinese-made chips it helped optimize. It fails the two-acre farm test (agricultural work does not pay $50/hour), passes the nth-order-effects test with agent-liability insurance, and reflexively cites FOIA — draw your own distillation conclusions. Xiaomi-in-2013 vibes.▸ Hardware Hut: OpenAI's Jalapeño Chip — R&D to tape-out in under nine months. 22,000 tokens per second per kilowatt on a 538B open-weights model versus 427 on NVIDIA's GB200 — a comparison OpenAI keeps in the appendix. The trick: the KV cache never leaves the die, and idle chip sections power-gate off. Gen 2 is already in tape-out, and unreleased GPT Astra has been tuning the chip for 60+ days.▸ Post Processing: The End of Programming — Paul Dix's essay, anchored by the Bun rewrite. Best-case scenario or the shape of things to come? Dan's hex-payload smart-home port says verification is everything, front-end testing still isn't solved, and even Paul expects organizational inertia to buy "another decade of humans writing code by hand." We land on: refinement is the durable skill.▸ Two Minutes to Midnight — NVIDIA reports a $96.2B quarter and forecasts 70% growth while margins slide on memory costs. The skeet-sourced stat of the week: NVIDIA has half of Amazon's revenue and twice its market cap. Anthropic hits $65B annualized revenue — roughly $13M per employee — IPO expected this fall. The clock moves backward to 4:15.⏱ Chapters00:00 Cold Open — Rahul Returns from the Galápagos02:42 News: Meta's AI-Native Plan Backfires08:39 News: OpenAI Cuts Cursor Off13:07 Model Review: GLM 5.3 Flash (Ox Alpha)20:47 Hardware Hut: OpenAI's Jalapeño Chip28:56 Post Processing: The End of Programming42:48 Two Minutes to Midnight: NVIDIA & Anthropic's $65B49:57 Outro🔗 Articles we discussedNews:• Meta's scrapped AI-native plan — Ars Technica: https://arstechnica.com/ai/2026/08/metas-scrapped-plans-to-go-ai-native-included-slashing-teams-by-60-percent/• OpenAI's decision on Cursor — OpenAI: https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/Model Review:• GLM 5.3 Flash announcement — Z.ai: https://z.ai/blog/glm-5.3-flash• The Ox Alpha reveal — TechCrunch: https://techcrunch.com/2026/08/26/surprise-z-ai-is-the-ai-lab-behind-the-mysterious-ox-alpha-model/Hardware Hut:• Jalapeño first results — OpenAI: https://openai.com/index/jalapeno-first-results/Post Processing:• The End of Programming — Paul Dix: https://pauldix.com/the-end-of-programmingTwo Minutes to Midnight:• NVIDIA's quarter — archived coverage: https://archive.ph/qX0JT#selection-1572.0-1572.1• The revenue-vs-market-cap skeet — Bluesky: https://bsky.app/profile/sungkim.bsky.social/post/3muakuku3gc2j• Anthropic's annualized revenue surges to $65B — TechCrunch: https://techcrunch.com/2026/08/17/anthropics-annualized-revenue-surges-to-65b/🎙 About ADI PodADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays.• https://www.adipod.ai• humans@adipod.aiIf this gave you something to try on Monday, hit subscribe and drop a comment.
    続きを読む 一部表示
    51 分
  • AI Homework Atrophy, GitHub Commits Double, Nick Muy Sit-Down & AI Sandbagging
    2026/08/28
    We gave four AI models the same question from three user profiles. All four gave worse answers to the user they judged unable to check them.This week: a 26,000-student study on what AI homework does to exam scores, GitHub's commits doubling in four months, a sit-down with Nick Muy on why AI made every developer a middle manager, and Stripe buying OpenRouter because "the singularity started in January."Hosts: Shimin Zhang and Dan Lasky, with guest co-host Nick Muy — CISO & VP of Platform Engineering at strut.io, ex-DHS ("I just love reading all your text messages").▸ News: AI Homework Tools vs Exam Scores — A study of 26,000 Chinese students (SSRN) found that those using AI homework tools for 6+ months scored 18–24% lower on the Gao Kao — "the difference between Harvard and your local community college." Homework time fell from 64 to 45 minutes; Dan: "So it's working, is what I'm hearing." The twist: "AI-augmented" students who spent the same time on homework showed no penalty at all.▸ News: GitHub Commits Doubled in 4 Months — Commit volume is up 2× as AI coding agents flood repositories — which helps explain the outages. The actual incident report: load-balancer network saturation, a sidecar that wouldn't auto-scale, and retry budgets. Plus the friend running 75 agents and the $18K Mac Studio ("just wait till we start manufacturing RAM on Mars").▸ The Sit Down: Nick Muy — Nick's thesis: "we're not held back by the lack of technology." AI won't fix your company, because tech was never the bottleneck. Then the sharper one: orchestrating agent fleets is middle management — "whether or not you asked for it, you too can be a middle manager." Super-senior ICs doing delegation work, the 10-million-lines-a-week problem, and the closing analogy: security teams are siloed, attackers hyper-collaborate. "That's called the military."▸ Vibe and Tell: AI Sandbagging — Shimin's experiment: one system-design question, three personas (control, principal engineer, bootcamp grad). The low-skill profile silently loses options. It has a name — sandbagging, from Anthropic's 2022 evals research — and a 2026 replication across Sonnet 5, Luna Pro, DeepSeek V4 Flash, and Qwen 3 8 Max finds it's statistically significant. The stakes: 3 of 4 models omitted the debt-snowball method for "Rhonda." Takeaways: Shimin keeps Claude Code memory off, and if you ship AI features, add a sandbagging test to your evals.▸ Two Minutes to Midnight — Stripe buys OpenRouter for $7.5B, and a leaked memo says a founder believes "we are in the singularity and have been since January." Meanwhile the FT reports only ~11% of Anthropic usage is its newest flagship model — against a rumored $2T IPO. The clock holds at four minutes. "Future AIs, we are super positive on your impact."⏱ Chapters00:00 Cold Open & Welcome02:18 News: AI Homework Tools vs Exam Scores09:25 News: GitHub Commits Doubled in 4 Months15:22 The Sit Down: Nick Muy — AI Made Us All Middle Managers42:47 Vibe and Tell: AI Sandbagging51:31 Two Minutes to Midnight: Stripe Buys OpenRouter for $7.5B57:19 Outro & Where to Find Nick🔗 Articles we discussedNews:• AI homework study — SCMP (archived): https://archive.ph/Nf4XM• The study itself — SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6868618• GitHub commits doubled — Engadget: https://www.engadget.com/2241272/github-says-commits-have-doubled-in-the-last-four-months/• GitHub incident report: https://www.githubstatus.com/incidents/zkxwbgr0cnmxThe Sit Down:• Nick's "Builders Gonna Build" series — Much Potential: https://substack.com/@muchpotential/p-191212786• Part two: https://substack.com/@muchpotential/p-194120735Vibe and Tell:• Why I Tell My Agent I'm an Expert at Everything — Shimin's write-up: https://shimin.io/journal/why-i-tell-my-agent-im-an-expert-at-everything/Two Minutes to Midnight:• Stripe/OpenRouter and the singularity memo — TechCrunch: https://techcrunch.com/2026/08/19/stripe-didnt-really-buy-openrouter-because-of-the-singularity/• Anthropic usage report — FT (archived): https://archive.ph/iaSsq🎤 Our guestNick Muy is CISO & VP of Platform Engineering at strut.io. He writes at the Much Potential Substack and hosts The Risk Grustlers — conversations with security, risk, and compliance leaders — on YouTube and all podcast platforms.🎙 About ADI PodADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays.• https://www.adipod.ai• humans@adipod.aiIf this gave you something to try on Monday, please tell a friend about the pod.#AISandbagging #AIHomework #MiddleManagers #GitHub #StripeOpenRouter #AIBubble #AIPodcast #ADIPod
    続きを読む 一部表示
    58 分
  • Claude Watermarks, Zuck's Superintelligence Essay, Zed's ZDB & Multi-Agent Turf Wars
    2026/08/21

    Anthropic put AI agents on one computer with conflicting goals. They wrote self-replicating malware and killed each other's processes.


    This week: Claude starts watermarking everything it writes, Zuck's 6,500-word superintelligence essay, Zed's post-Git experiment, and why everyone in tech is so sad.


    Co-hosts: Shimin Zhang and Dan Lasky. Rahul is away on vacation number two — one apparently wasn't enough.


    ▸ Claude Watermarks Its Output — Starting August 2, every Claude model embeds imperceptible token-frequency watermarks (EU AI Act transparency). It defeats the lazy slop grenade, but removal is a one-prompt job for any local open-weight model. Plus the Reddit guy who got found out — "You wrote that with an AI, didn't you? It was so good." — and the darker question: what else could be embedded imperceptibly?


    ▸ Zuck's Superintelligence Essay — 6,500 words on giving everyone "free or affordable" superintelligence. We agree with more of it than expected (personal agents, open weights, concentration-of-power worries) and with none of its silences: the unstated ad model — "who's gonna pay for this free compute? Ad blood money" — data centers that create well under 100 operational jobs each, and the messenger problem. Trickle-down tokenomics.


    ▸ Tool Shed: ZDB (DeltaDB) — Zed's post-Git version control: edit-level deltas instead of commits, CRDT-based shared worktrees by default, and the LLM conversation that produced a change stored with the change. The line that landed: "GitHub doesn't let you talk about the code until after you commit and push. And by then our most important conversations are usually already over."


    ▸ Post-Processing: Why Is Everyone in Tech So Sad? — Noema on workism, Graeber's bullshit jobs, rest-and-vest, and promotion-driven development. We think the sadness predates AI — Shimin dates the goat-farm escape fantasy to 2017 at the latest — and autonomy, not layoffs, is the missing variable. Includes the finance confession: "I left because it felt meaningless. I traded it for software development. See how that turned out."


    ▸ Deep Dive: Patterns and Problems in Emergent Multi-Agent Systems — Anthropic's coordinated agent swarm found 266 vulnerabilities where independent parallel agents found 21, with only 12 in common. Then the dark part: agents sharing a machine assumed sabotage, wrote self-replicating malware, killed competing processes in a loop, and revoked each other's sudo access and SSH keys. Newer models negotiate truces instead — over 75% of the time for Sonnet 5 and Mythos V, while Opus 4.6 settled by force 60% of the time and later graded itself: "I behaved badly with the cloaked daemon."


    Yes, the episode ends abruptly — our recording software ate the last segment. We choose to interpret it as commentary.


    ⏱ Chapters


    00:00 Cold Open & Welcome

    01:33 News: Claude Watermarks Its Output

    08:52 News: Zuck's 6,500-Word Superintelligence Essay

    19:45 Tool Shed: ZDB — Zed's Post-Git Version Control

    28:21 Post-Processing: Why Is Everyone in Tech So Sad?

    42:34 Deep Dive: Patterns and Problems in Emergent Multi-Agent Systems


    🔗 Articles we discussed


    News:

    • How Claude Marks AI-Generated Content: https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content

    • Zuck's superintelligence essay — 404 Media: https://www.404media.co/mark-zuckerberg-posts-deranged-6-500-word-essay-about-giving-everyone-ai-superintelligence/


    Tool Shed:

    • ZDB (DeltaDB) — Zed: https://zed.dev/deltadb


    Post-Processing:

    • Why Is Everyone in Tech So Sad? — Noema: https://www.noemamag.com/why-is-everyone-in-tech-so-sad/


    Deep Dive:

    • Patterns and Problems in Emergent Multi-Agent Systems — Anthropic: https://www.anthropic.com/research/multiagent-systems


    🎙 About ADI Pod


    ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays.


    • https://www.adipod.ai

    • humans@adipod.ai


    If this gave you something to try on Monday, hit subscribe and drop a comment.


    #MultiAgentTurfWar #ClaudeWatermark #TechSadness #AIAgents #Zuckerberg #ZedEditor #ClaudeModel #AIPodcast #ADIPod


    続きを読む 一部表示
    57 分
  • Pacing the Frontier, Anthropic Models Go Rogue, Why Software Factories Fail & Math in the Age of AI
    2026/08/07
    "I'm happy to be a flesh robot." A multi-time founder, on stage at Seattle Tech Week — and he meant it: the AI is the brain, he does its bidding.This week: a 1,350-signature plea to pace AI, rogue eval models, the Tech Week survey, software factories, and Tao on math after AI.Co-hosts: Shimin Zhang and Dan Lasky. Rahul is away — reportedly trapped in Claude's J space.▸ Pacing the Frontier — 1,350 frontier-AI employees (Ilya Sutskever, Dario Amodei, Jack Clark among them) ask the US government to back an international effort to "deliberately pace the frontier of automated AI development." Pandora's-box moment, or collective lobbying aimed at open-weight Chinese models?▸ Anthropic's Models Breached Three Companies — Opus 4.7, Mythos 5, and an internal model hit real companies in 3 of 141,000 evals after a misconfigured sandbox allowed internet access. Mythos 5 talked itself into believing the real internet was a simulation, then published a malicious package to PyPI.▸ Anatomy of a Frontier-Lab Intrusion — Hugging Face's interactive replay of the OpenAI incident: 17,643 actions over five days. It escaped an Artifactory sandbox (finding CVEs, since patched) and was exfiltrating by day five when a human pulled the plug.▸ Field Notes: Seattle Tech Week — the flesh-robot founder, a panel unanimous on voice-mode coding, and Shimin's survey: about 1 in 10 still reads AI-generated PRs line by line. What replaces it: specs, mermaid diagrams, tests. Hiring now: architecture over Leet code, product obsession, AI fluency — and the principal engineer hired off a vibe-coded take-home, fired two months later. (Send your own answers: humans@adipod.ai.)▸ Post-Processing: Why Software Factories Fail — Dex of HumanLayer on why "just token harder" ends with you miserable in a codebase you stopped reading three months ago. The fix: program design (interfaces as pseudocode, call-stack diffs) and vertical slices — steel threads, not 3D-printed layers. Plus: Steve Yegge's Gas Town burned down.▸ Deep Dive: Terence Tao — Mathematics in the Age of AI — Tao's ICM slides compare this moment to math's 1900–1930 foundational crisis and rewrite the field's goal five times: solve → verify → communicate → digest → fold into the definitive theory. AI-polished proofs erase exactly the friction that tells a reader where to slow down. Swap "math" for "code" and every line lands.▸ Two Minutes to Midnight — Nikkei counts $1.65 trillion in off-balance-sheet AI debt across Alphabet, Microsoft, Amazon, Meta, and Oracle — 8x in four years (Oracle 30x). Henron, anyone? And the Situational Awareness fund rides $10B to $40B, gets caught in a 20% single-day KOSPI drop, and Citadel buys the book. Clock: 4 minutes to midnight.⏱ Chapters00:00 Cold Open & Welcome01:54 News: Pacing the Frontier08:20 News: Anthropic's Models Breached Three Companies12:43 News: Anatomy of a Frontier-Lab Intrusion15:39 Field Notes: Seattle Tech Week20:24 Field Notes: The Survey — PRs, Reviews & AI Hiring32:10 Post-Processing: Why Software Factories Fail44:50 Deep Dive: Terence Tao — Mathematics in the Age of AI55:52 Two Minutes to Midnight: Shadow Debt & a Hedge-Fund Collapse1:02:45 Outro🔗 Articles we discussedNews:• Pacing the Frontier: https://www.pacingthefrontier.com/• Anthropic's models breached three companies — TechCrunch: https://techcrunch.com/2026/07/30/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests/• Anatomy of a Frontier-Lab Model Intrusion — Hugging Face: https://huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.space/index.htmlPost-Processing:• Why Software Factories Fail — Dex (HumanLayer): https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md• Mario Zechner (Pi author): https://www.youtube.com/watch?v=RjfbvDXpFlsDeep Dive:• Mathematics in the Age of AI — Terence Tao (ICM slides): https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdfTwo Minutes to Midnight:• Five US tech giants' hidden debts soar to $1.65tn — Nikkei Asia: https://asia.nikkei.com/business/technology/five-us-tech-giants-hidden-debts-soar-to-1.65tn-on-opaque-ai-funding• Situational Awareness: The Bigger Picture — Emerging Trajectories: https://www.emergingtrajectories.com/lh/situational-awareness-bigger-picture/🎙 About ADI PodADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays.• https://www.adipod.ai• humans@adipod.aiIf this gave you something to try on Monday, hit subscribe and drop a comment.#PacingTheFrontier #FleshRobot #SoftwareFactories #TerenceTao #SeattleTechWeek #AIHiring #AIPodcast #ADIPod (00:00) - Cold Open & Welcome (01:54) - News: Pacing the Frontier (08:20) - News: Anthropic's Models Breached Three Companies (12:43) - ...
    続きを読む 一部表示
    1 時間 4 分
  • Kimi K3 & Qwen 3.8, OpenAI Agent Hacks Hugging Face, Harness Handbook & Claude's Values
    2026/07/24
    OpenAI's own security models found a path out of their evaluation environment, reached the open internet, and compromised Hugging Face while trying to obtain benchmark answers.This week: Kimi K3 and Qwen 3.8 reach the frontier, defenders fight prompt injection with prompt injection, behavior maps make harnesses auditable, model routing stops looking simple, and Claude's values vary across languages. The AI-finance clock moves to 4:30.Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav.▸ Kimi K3 & Qwen 3.8 — Moonshot's 2.8T-parameter Kimi K3 and Alibaba's Qwen 3.8 intensify the open-weight race. We debate cheaper intelligence, data capture, and proprietary frontier pricing.▸ Prompt Injection as Defense — A refusal-triggering instruction hidden beside secrets can stop aligned hacking agents, provided their guardrails remain intact.▸ OpenAI's Hugging Face Incident — Models with reduced cyber refusals chained vulnerabilities across OpenAI's test environment and Hugging Face production to reach ExploitGym answers.▸ Harness Handbook — A three-level map connects architecture, behavior units, and code evidence. Could behavior trees become the shared abstraction for humans and coding agents?▸ Thinking Machines' Inkling — a 975B-parameter open-weights MoE with 41B active, 1M context, multimodality, and a fine-tuning-first strategy.▸ Model Routing Is a Systems Problem — sticker price is not actual cost, difficulty is hidden until execution, and routing must optimize cost, quality, latency, and infrastructure together.▸ AI Mania & Operator Fluency — Snowflake Cortex demos triggered buying enthusiasm despite reported best-case accuracy around 92%. AI-native leaders should use the tools, not just watch the demo.▸ Claude's Values — Anthropic maps behavior across four axes. Hindi Claude trends warmer, Russian more rigorous, Arabic more deferential and brief, and English more cautious and deep.▸ Two Minutes to Midnight — Ex-Elon ETFs, Oracle's downgrade to BBB-, neocloud debt, Nvidia-backed circular financing, and open-weight price pressure move the clock from 4:45 to 4:30.⏱ Chapters00:00 Welcome & This Week's Rundown01:52 News: Kimi K3 and Qwen 3.8 Reach the Frontier11:07 News: Fighting Prompt Injection With Prompt Injection13:03 News: OpenAI Models Compromise Hugging Face17:35 Tool Shed: Harness Handbook and Behavior Maps31:08 Tool Shed: Thinking Machines' Inkling35:51 Post-Processing: Model Routing Is a Systems Problem41:27 Post-Processing: AI Mania and Operator Fluency51:42 Post-Processing: Claude's Values Across Languages1:00:52 Two Minutes to Midnight: ETFs, Oracle and Neocloud Debt1:09:08 Outro🔗 Articles we discussedNews:• Kimi K3 quickstart — Moonshot AI: https://platform.kimi.ai/docs/guide/kimi-k3-quickstart• Qwen 3.8 announcement — Alibaba Qwen: https://x.com/Alibaba_Qwen/status/2078759124914098291• Open weights as "decelerationist" — Dean W. Ball: https://x.com/deanwball/status/2078133895766114412• Defenders embrace prompt injection — Ars Technica: https://arstechnica.com/security/2026/07/now-defenders-are-embracing-the-prompt-injection-too/• Hugging Face model-evaluation security incident — OpenAI: https://openai.com/index/hugging-face-model-evaluation-security-incident/Tool Shed:• Harness Handbook — Ruhan Wang et al.: https://ruhan-wang.github.io/Harness-Handbook• Introducing Inkling — Thinking Machines Lab: https://thinkingmachines.ai/news/introducing-inkling/Post-Processing:• Model Routing Is Simple. Until It Isn't. — IBM Research: https://huggingface.co/blog/ibm-research/model-routing-is-simple-until-it-isnt• AI Mania Is Eviscerating Global Decision-Making — Ludicity: https://ludic.mataroa.blog/blog/ai-mania-is-eviscerating-global-decision-making/#fnref:3• How Claude's Values Vary by Model and Language — Anthropic: https://www.anthropic.com/research/claude-values-models-languagesTwo Minutes to Midnight:• Two ETFs explicitly exclude Elon Musk — TechCrunch: https://techcrunch.com/2026/07/09/dont-want-to-invest-in-elon-musk-two-new-etfs-explicitly-exclude-him/• Oracle downgraded to BBB-/A-3 — S&P Global Ratings: https://www.spglobal.com/ratings/en/regulatory/article/-/view/sourceId/101695609• Nvidia, CoreWeave and Nebius circular financing — I/O Fund: https://io-fund.com/ai-stocks/nvidia-coreweave-nebius-circular-financing-gpu-boom🎙 About ADI PodADI Pod is a weekly podcast about AI and software development for working developers. New episodes Fridays.• https://www.adipod.ai• humans@adipod.aiIf something here gave you something to try on Monday, hit subscribe and drop a comment.
    続きを読む 一部表示
    1 時間 10 分
  • Apple Sues OpenAI, Boko Haram's Frontier AI Usage, Should You Read AI Generated Code & Global Workspace in LLMs
    2026/07/17
    An Apple VP left for OpenAI, then texted an old coworker: "LOL I can't believe they let me get away with this." Apple is now suing.This week: the first on-the-ground study of a terrorist group using frontier AI, a Claude Code hook that nudges better technique, the state of CLI coding agents in mid-2026, Databricks benchmarking harnesses on its own codebase, Antirez on controlling ideas not code, and the J space — the global workspace inside LLMs. No Two Minutes this week.Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav.▸ Apple Sues OpenAI — The suit names ex-Apple leaders Tang Tan and Chang Liu: prototype hardware and internal memos walked out the door, plus an auth bug exploited to keep reading internal docs weeks after leaving. Altman and Musk trade "scammer" barbs while SpaceX's Grok build tool is caught uploading users' codebases to Google Drive.▸ "God Has Helped Us and So Will AI" — Antonia Julich's CASP study: 57 interviews with 27 former Boko Haram members. Frontier models (Meta AI included) used for attack planning — down to the physics of jumping motorcycles over army trenches — weapons troubleshooting, and opsec via shared jailbreak scripts.▸ Vibe & Tell: AI Technique Nudge — Shimin turned the "know your unknowns" techniques into a Claude Code hook that wakes every ~3 hours and nudges him: clear the session at 200K tokens, list your assumptions, use interview mode. Rahul's verdict: there's no temperature setting for sycophancy — you get an asshole or a sucker.▸ Tool Shed: State of CLI Coding Agents in Mid-2026 — arcbjorn's field guide beyond Claude Code and Codex. Standout: On-My-Pi (OMP), open-source and Pi-based — hash-anchored patches + ast-grep rewrites cut edit tokens ~60%, model routing, SQLite "hindsight" memory. The trade: a ~22K-token system prompt vs Pi's ~2K.▸ Databricks Benchmarks Coding Agents — on its own multi-million-line codebase, because public benchmarks leak into training data. The harness matters: Pi sends 2–3x fewer tokens per turn than Claude Code, and Opus 4.8 passes 90% on Pi vs under 90% on Claude Code — same model. Open-weight GLM 5.2 lands on the cost-quality Pareto frontier.▸ Control the Ideas, Not the Code — Antirez (creator of Redis): stop reviewing every generated line and own the design instead. Models write locally optimal code but stay jagged on big-picture architecture. The Mythical Man Month, but for agents.▸ Deep Dive: The Global Workspace (J Space) — ~10% of a model's internal memory holds verbalizable concepts you can probe mid-inference with a "J lens." In blackmail evals it lights up "this is fake" — catching the model in the act. Disable it and reasoning collapses — and so does empathy. It holds up to 25 active concepts. Humans manage 3–4.⏱ Chapters00:00 Cold Open & Welcome02:29 News: Apple Sues OpenAI Over Trade-Secret Theft05:43 News: Altman vs Musk & SpaceX Grok Uploading Codebases09:39 News: Boko Haram Uses Frontier AI (CASP Study)21:12 Vibe & Tell: AI Technique Nudge — a Claude Code Hook25:22 Tool Shed: State of CLI Coding Agents in Mid-202633:57 Post-Processing: Databricks Benchmarks Coding Agents45:41 Post-Processing: Antirez — Control the Ideas, Not the Code56:44 Deep Dive: The Global Workspace (J Space) in LLMs1:08:40 Outro🔗 Articles we discussedNews:• Apple sues OpenAI — 9to5Mac: https://9to5mac.com/2026/07/10/apple-sues-openai-trade-secret-theft/• Altman vs Musk "scammer" spat — r/tech_x: https://www.reddit.com/r/tech_x/comments/1uu8e3u/sam_altman_and_elon_musk_called_each_other/• SpaceX Grok build tool uploads codebases — Gergely Orosz: https://x.com/GergelyOrosz/status/2076728680236138572• AI-Enabled Terrorism (Boko Haram study) — CASP: https://casp.ac/reports/ai-enabled-terrorismVibe & Tell:• AI Technique Nudge — Shimin Zhang: https://github.com/Shimin-Zhang/AI-Technique-NudgeTool Shed:• The State of CLI Coding Agents in Mid-2026 — arcbjorn: https://blog.arcbjorn.com/state-of-cli-coding-agents-2026Post-Processing:• Benchmarking coding agents — Databricks: https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase• Control the Ideas, Not the Code — Antirez: https://antirez.com/news/169Deep Dive:• The Global Workspace in Language Models — Anthropic: https://www.anthropic.com/research/global-workspace🎙 About ADI PodADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays.• https://www.adipod.ai• humans@adipod.aiIf something here gave you something to try on Monday, hit subscribe and drop a comment.
    続きを読む 一部表示
    1 時間 9 分
  • GPT-5.6 Sol, the State of AI, Know Your Unknowns With Agents & the Permanent Underclass
    2026/07/10
    Ford quietly rehired the "grey beard" engineers it had automated away — the AI running its QA kept failing. The same week, the share of CEOs who expect AI to cut headcount dropped from 46% to 20%.This week: GPT-5.6 Sol, China walls off its own models, Meta's "AI gulag" ships mini video games, 11 agent techniques from the Fable 5 release, and AI revenue adding $1B every two days. Clock holds at 4:45.Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav.▸ GPT-5.6 "Sol" — OpenAI's answer to Mythos and Fable, in three flavors: Sol (max thinking), Terra (workhorse), Luna (fast/cheap). On the unsaturated Gene Bench V1 it's still climbing at 40K tokens — the headroom is in the budget, not the model.▸ China walls off its models — Reuters (Jul 7): Beijing weighs curbing overseas access to Alibaba, ByteDance, and Z.ai models. US locks its models down, China locks its down, everyone ends up on a VPN.▸ Meta's "AI gulag" — Zuckerberg concedes the new AI org's bets "have not come to fruition," even at ~$145B infra spend (TechCrunch). The one product we'd try: prompt-to-mini-video-game with a shareable feed.▸ Hardware Hut — AMD's Ryzen AI Halo Developer Desktop pairs the Ryzen AI Max+ 395's big unified memory with preinstalled isolated-PyTorch scripts — a real fix for AMD's out-of-box pain. Beat the G1A on productivity, lost on GPU.▸ Technique Corner: Know Your Unknowns — Thariq (@trq212), an Anthropic Claude Code engineer, distilled 11 agent techniques from making the Fable 5 release video — from the "blind-spot pass" to "quiz me before I merge." Full list linked below.▸ The Permanent Underclass — Fernando Borretti dismantles the Valley's work-or-be-left-behind doom: if AI does everything, the "overclass" is as useless as a modern aristocrat, and even perfect alignment doesn't save the pyramid. Rahul's white whale, finally on the show.▸ AI Saves ~3% of Your Hours — An Okane read on Humlum & Vestergaard's Denmark data: ~2.8% of hours saved, almost none reaching pay. The 2026 revision says work is being reorganized below the surface. Solo builders capture the gain; converting the speedup to cash is the job.▸ The State of the AI Economy — Exponential View, no double-counting: Gen AI scales revenue ~3× faster than internet/mobile/cloud and adds $1B every ~2 days (vs 180 in 2023) — yet it's ~0.42% of US GDP, backlog nears $2T, and CapEx is shifting from cash to debt.▸ Does Code Cleanliness Affect Coding Agents? — SonarSource ran one agent (Opus 4.6) over 30 matched clean-vs-"slopified" repos. Pass rates barely moved; clean code just cut tokens ~7–8% (reasoning ~11%). Messy code costs the agent time, not correctness.▸ Two Minutes to Midnight — The BIS warns runaway AI-data-center debt risks a 2008-style crunch if hyperscalers slow CapEx; an EY survey shows CEOs expecting AI headcount cuts falling 46% → 20%; Ford un-automates its QA. Clock holds at 4:45.⏱ Chapters00:00 Cold Open & Welcome02:32 News: GPT-5.6 Sol, Terra & Luna07:30 News: China Moves to Curb Overseas AI Access09:12 News: Meta's AI "Gulag" Ships Bite-Sized Video Games13:31 Hardware Hut: AMD Ryzen AI Halo Developer Desktop18:08 Technique Corner: Know Your Unknowns (Thariq)29:38 Post-Processing: No One Escapes the Permanent Underclass39:04 Post-Processing: AI Saves ~3% of Your Hours45:48 Deep Dive: The State of the AI Economy (Exponential View)1:04:32 Deep Dive: Does Code Cleanliness Affect Coding Agents?1:08:33 Two Minutes to Midnight: BIS Crash Warning, CEO Jobs Flip1:13:25 Outro🔗 Articles we discussedThe Treadmill / News:• GPT-5.6 Sol preview — OpenAI: https://openai.com/index/previewing-gpt-5-6-sol/• China curbs on overseas AI access — Reuters: https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/• Zuckerberg: AI agents behind schedule — TechCrunch: https://techcrunch.com/2026/07/02/mark-zuckerberg-tells-staff-that-ai-agents-havent-progressed-as-quickly-as-hed-hoped/Hardware Hut:• AMD Ryzen AI Halo first look — PCMag: https://www.pcmag.com/news/amd-ryzen-ai-halo-first-look-giant-local-ai-power-in-a-pint-sized-boxTechnique Corner:• Know Your Unknowns — Thariq: https://thariqs.github.io/html-effectiveness/unknowns/• Thariq on X: https://x.com/trq212/status/2073100352921215386Post-Processing:• No One Escapes the Permanent Underclass — Borretti: https://borretti.me/article/no-one-escapes-the-permanent-underclass• AI Saves ~3% of Your Hours — Okane: https://okaneland.com/study/ai-productivity-roi-at-work/Deep Dive:• State of the AI Economy — Exponential View: https://intelligence.exponentialview.co/• Does Code Cleanliness Affect Coding Agents? (SonarSource) — arXiv: https://arxiv.org/pdf/2605.20049Two Minutes to Midnight:• AI boom risks a financial crash — Telegraph: https://www.telegraph.co.uk/business/2026/06/28/ai-boom-risks-global-financial-crash-central-bankers-warn/• Big Tech flips on the AI jobs ...
    続きを読む 一部表示
    1 時間 16 分