Demo green, prod wrong: failure types founders should name
カートのアイテムが多すぎます
カートに追加できませんでした。
ウィッシュリストに追加できませんでした。
ほしい物リストの削除に失敗しました。
ポッドキャストのフォローに失敗しました
ポッドキャストのフォロー解除に失敗しました
-
ナレーター:
-
著者:
Hey! I'd love to hear your thoughts, send me a voice note.
Sources:
• OWASP Top 10 — https://owasp.org/www-project-top-10-for-large-language-model-applications/
• OWASP Prompt Injection — https://cheatsheetseries.owasp.org/cheatsheets/LLM_Prompt_Injection_Prevention_Cheat_Sheet.html
• Check Point PuzzleMask — https://research.checkpoint.com/2026/puzzlemask-abusing-plain-prose-as-a-covert-ai-attack-vector/
• Invariant Labs MCP poisoning — https://invariantlabs.ai/blog/mcp-security-notification-tool-poisoning-attacks
• Embrace The Red SpAIware — https://embracethered.com/blog/posts/2024/chatgpt-macos-app-persistent-data-exfiltration/
• PoisonedRAG — https://www.usenix.org/conference/usenixsecurity25/presentation/zou-poisonedrag
• Authorization-First Retrieval — https://aclanthology.org/2026.trustnlp-main.15/
• Pillar Security Deadbugz — https://www.pillar.security/blog/deadbugz-currently-active-mcp-supply-chain-campaign
• APIsec Labs A2A peers — https://labs.apisec.ai/research/articles/hijacking-google-adk-malicious-a2a-peers/
• OpenAI Hugging Face report — https://openai.com/index/hugging-face-incident-and-the-road-ahead/
• Wiz Off Guard — https://www.wiz.io/blog/off-guard-breaking-litellm-from-authentication-bypass-to-cloud-compromise
• CISA KEV alert — https://www.cisa.gov/news-events/alerts/2026/09/02/cisa-adds-seven-known-exploited-vulnerabilities-catalog
• Anthropic Threat Intelligence — https://www.anthropic.com/threat-intelligence-report-september-2026
• Anthropic cybersecurity evals — https://www.anthropic.com/research/investigating-incidents-cybersecurity-evals
• Stochasticity in Agentic Evaluations — https://arxiv.org/html/2512.06710v1
• OSWorld-Verified — https://xlang.ai/blog/osworld-verified
Chapters:
00:00:00 Intro music
00:00:06 Cold open
00:01:44 Definitions
00:03:43 Why agents need a different map
00:04:29 Family 1: Injection
00:06:42 Family 2: Tool abuse
00:08:46 Family 3: Context and memory
00:10:24 Family 4: RAG and vector stores
00:12:21 Family 5: Connectors and MCP
00:13:54 Family 6: Multi-agent handoffs
00:15:51 Family 7: Auth gaps
00:17:27 Family 8: Exfiltration via tools
00:18:49 Family 9: Fake complete
00:20:14 Family 10: Nondeterminism
00:21:17 Ship gate for a small team
00:23:47 Homework
00:24:57 Close