E09: How to Talk AI Agents to Find Sneaky Production Bugs
カートのアイテムが多すぎます
カートに追加できませんでした。
ウィッシュリストに追加できませんでした。
ほしい物リストの削除に失敗しました。
ポッドキャストのフォローに失敗しました
ポッドキャストのフォロー解除に失敗しました
-
ナレーター:
-
著者:
Host Ran Aroussi (Old School / New Tech) shares a recent “war story” from his agency Automaze: a hard-to-reproduce mobile bug where calls were sometimes dropped only on a client’s devices, which the team couldn’t reproduce for 2–3 weeks. After two long sessions using a Droid Factory setup with agents and a “small council” debate between models (e.g., Fable and Sol), he got the team unstuck, found the issue, and then turned the session logs into an internal guide and public article on his methodology. Key practices include distrusting agent conclusions (treat “preexisting/flaky/unrelated” as hypotheses), demanding proof (including his Proof library requiring video evidence), using extensive unit and end-to-end tests, asking for a numeric confidence level before production, controlling environments, handling merge conflicts with full review/testing, switching models for independent review, and re-spec’ing from scratch to detect drift from the implementation plan.