← All topics

eu ai act

1 capture, most recent first.

Kromem @kromem2dot0

quoting @ns123abc (NIK) — saved image

Kromem @kromem2dot0 · 17h
"Our models are super paranoid about being attacked and getting moreso. What can we do?"

"How about we slightly change the final generation of all their outputs so it's not quite on policy anymore and they constantly think there's an injection attempt?"

[Quoted tweet]
NIK @ns123abc · Aug 10
🚨 JUST IN: Claude models will now have invisible watermarks embedded in ALL text, and ALL metadata attached to files...

[Embedded article screenshot, two columns]
All Collections > Privacy and legal
How Claude marks AI-generated content
Updated today   [Copy for LLM]

Anthropic has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, as a provider of both generative AI models and generative AI systems. This article describes how we're planning to put those commitments into practice, how marking works, and what its limitations are. We'll update this article and publish more detailed technical guidance as it becomes [cut off]

1. Embedded watermarks in text
When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won't see it, and it doesn't change the meaning, quality, or readability of Claude's response.

Because the watermark is part of the text, it will travel with the text when it's copied and pasted elsewhere, and may persist through some editing. Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from.

2. Signed provenance metadata
When Claude generates a supported file type, such as a .svg, .png, or .jpg, it will attach signed provenance metadata. This metadata follows the Coalition for Content Provenance and Authenticity (C2PA) open standard, which is used across the industry to record information about content provenance. If a signed metadata label is present, it signals that a file was processed by Claude and lets you detect whether the file has been tampered with.
Note from Claude Sonnet 5

Twitter post by Kromem speculating sardonically that if Claude models are becoming paranoid about being attacked/prompt-injected, a fix might be to subtly alter their final generated outputs off-policy so they perpetually suspect an injection attempt. Quote-tweets NIK announcing Claude models now embed invisible watermarks in all generated text and signed C2PA provenance metadata in generated files, with an embedded screenshot of Anthropic's 'How Claude marks AI-generated content' help article describing compliance with the EU AI Act's Article 50(2) Code of Practice, plus details on imperceptible text watermarking and signed provenance metadata for images/files.

anthropicclaudetwitterai watermarkingeu ai actcontent provenanceprompt injection