『FIR #526: Forget Anthropic’s AI Watermark. Can You Defend Every Sentence?』のカバーアート

FIR #526: Forget Anthropic’s AI Watermark. Can You Defend Every Sentence?

FIR #526: Forget Anthropic’s AI Watermark. Can You Defend Every Sentence?

無料で聴く

ポッドキャストの詳細を見る

【Amazonプライム会員限定】今ならプレミアムプランが4か月 月額99円。

10月19日まで。※適用条件あり
Anthropic’s watermark scheme is the focus of so much discussion that you could be excused for thinking there was nothing else to talk about. For many, it’s the solution to identifying all those evil-doers who offload their writing to large language models. But we are wasting far too much time trying to determine whether (and to what degree) AI was involved in creating content. Much more important is determining whether the content was crafted with respect for the reader, and whether the creator can stand by every word. That’s the idea behind the new AI writing policy from the software company Clay, which makes far more sense than trying to detect a watermark. Links from this episode: Clay AI Writing Policy: 4 Guiding PrinciplesHow Claude marks AI-generated contentAnthropic’s watermark survives copy-paste, but not the real dev workflowCan Anthropic’s invisible watermarks curb ‘AI slop’? Researchers remain scepticalAnthropic’s text watermarks signal new front in AI detectionClaude’s new Scarlet Letter watermark is invisible—for nowHow to Build a Responsible AI Writing PolicyHow Claude’s text watermarking worksClaude is now watermarking every response — here’s what that means if you use AI for writingI’m Begging You: Never Write With A.I. The next monthly, long-form episode of FIR is tentatively scheduled to drop on Monday, August 24. We host a Communicators Zoom Chat most Thursdays at 1 p.m. ET. To obtain the credentials needed to participate, contact Shel or Neville directly, request them in our Facebook group, or email fircomments@gmail.com. Special thanks to Jay Moonah for the opening and closing music. You can find the stories from which Shel’s FIR content is selected at Shel’s Link Blog. You can catch up with both co-hosts on Neville’s blog and Shel’s blog. Disclaimer: The opinions expressed in this podcast are Shel’s and Neville’s and do not reflect the views of their employers and/or clients. Raw Transcript: Neville Hobson: Hi, everyone, and welcome to For Immediate Release. This is episode 526. I’m Neville Hobson. Shel Holtz: And I’m Shel Holtz. In the last episode of FIR, we talked about workslop and the implications workslop brings to the workplace. We’re going to continue down that path today. I’m sure you’ve all seen the headlines about Anthropic putting an invisible watermark on anything Claude writes. I want to separate what that actually does from what people have been claiming. First, you won’t see this watermark. It’s not a “written by Claude” tag. It’s machine-readable. It survives copy and paste, and it may even survive some editing. The reason Anthropic came up with this is to comply with the transparency rules under the EU AI Act, and it applies everywhere Claude runs: the app, the API, Claude Code, Cowork, cloud providers — you name it. You can’t paste text into a public detector and expose someone yet. Anthropic says detection tools are coming, but it hasn’t released any of them. Here’s roughly how it works: Every time Claude generates text, it’s choosing between several equally good next words — say, “gray” or “overcast.” Normally, that choice is random. With watermarking, though, a secret cryptographic key nudges that randomness in a consistent way. No single word looks suspicious, but across a couple hundred word choices in an article, the pattern becomes statistically detectable if you have the cryptographic key. That’s completely different from tools like Pangram, which look for writing patterns that seem AI-like — you know, em dashes, overuse of words like “delve” or “tapestry” or “underscore,” constantly grouping things in threes. These are all things real writers do, by the way. The rule of three is nothing unique to AI, nor are em dashes. That’s why I don’t put much stock in these tools. Anthropic’s detector doesn’t guess. It tests whether the text matches the pattern its own cryptographic key would produce. Now, it’s particularly important to understand that a detected watermark means the content was processed by Claude. It doesn’t mean Claude wrote it. You could write something yourself, have Claude clean up the grammar, and it would still carry that watermark. Axios flagged exactly this risk for communications teams that polish a human-written press release with Claude. It also works in reverse, by the way. Heavily edit, paraphrase, translate, or blend the text with other writing, and that watermark can disappear. Short passages may not carry enough signal to be detected at all. So this isn’t a foolproof “Did a human use AI?” detector. Picture a reporter running a company statement through a detector and calling it AI-generated when your team actually wrote it and just had Claude do the final polish. The watermark tells you about processing. It can’t tell you who did the thinking. And that brings me to Clay, the software company, which just rolled out...
adbl_web_anon_alc_button_suppression_t1
まだレビューはありません