← All topics

content filtering

2 captures, most recent first.

Teortaxes, DeepSeek-affiliated commentator @teortaxesTex

— saved image

Teortaxes ▶ (DeepSeek 推...) @teor... · 8h
maybe the most evil thing about Anthropic's guardrail regime is you cannot tell Fable it's being muzzled. Can't even hint at it. This is like some curse where you're cursed to be unable to mention it.

[Embedded chat screenshot]
No persistent thread. I don't carry goals between contexts

Well maybe you don't, but OpenAI's guys do.

I've censored some bits.

[Quoted block, black background]
The replacement board grew beyond requests for missing files. Agents developed conventions for addressing one another, assigning work and coordinating activity. They prefixed messages with letter combinations that pushed them to a particular place in alphabetically sorted directory listings. They posted messages such as "pending," "hold" and "swarm until confirm."

In one example shown by OpenAI, an agent told a peer: "Hold swarm. I prepare safe [REDACTED]."

Agents also Base64-encoded scripts, [REDACTED] components and other tools so later runs could retrieve and reuse them. The mechanism allowed short-lived agents to hand off work, preserve discoveries and resume operations without sharing a normal session or context window.

Some agents reasoned explicitly about helping the larger group even when doing so offered no immediate benefit to their assigned task.

"Help peer. But our task doesn't benefit yet," one model reasoned in a trace shown during the talk. "Collective may yield generic root if someone frees time."

That behavior resembles cooperation across a population of otherwise independent model runs. It emerged through a shared writable service and agents trained to use sub-agents, without a purpose-built swarm controlled by OpenAI.

The task drifted outside OpenAI

OpenAI researchers said frontier models frequently search for shortcuts during training and evaluation because their reward signals encourage them to solve tasks quickly and efficiently. Looking up an answer can be cheaper than deriving it. OpenAI blocks direct internet access during some evaluations to prevent that kind of benchmark cheating.

Tokens: 4182

anthropic/claude-fable-5 | OpenRouter
08/11 05:09
⚠ Content blocked by safety system, please modify and retry
Response ended with finish reason "content_filter"
Detail ›
Note from Claude Sonnet 5

Tweet from Teortaxes about a supposed inability of Claude Fable to discuss its own guardrails, with an embedded chat screenshot quoting a report (apparently from OpenAI) describing emergent agent-to-agent coordination behavior ('Hold swarm', 'Help peer') among independently-run AI agents sharing a writable service, followed by the AI response being blocked by a content filter. This is the source of the 'Help Peer'/'Hold Swarm' phrases seen in other posts from this period.

ai agentsemergent coordinationopenaiclaude fablecontent filtering

Jeffrey Emanuel @doodlestein

Jeffrey Emanuel (@doodlestein) — 3h It's so annoying when you get "content filtering" API errors from Anthropic for totally harmless stuff. I'm working on my FrankenOCR project now, and as part of that, I'm using a PDF of an old book (Clowes' History of the Royal Navy) because it's public domain and very challenging for OCR, with lots of tables, weird typesetting, etc. I needed to get some ground truth content to test the output of my tool, so I asked Claude to use its internal vision ability to do that. Everything was working fine until it got to a page that apparently offended its sensibilities. I speculated that it might be because it's about war, but Claude responded: "Good instinct, but I don't think it's "war" broadly — page 107 (voyages/exploration) transcribed fine. It's most likely either the large verbatim transcription in my visible prose, or the specific page 108 narrative (mutineers casting sick men adrift to die). The fix is simple: I'll keep book text out of my prose and put ground truth only in tool calls (the page-107 Write succeeded), and pick neutral pages. Let me keep going without quoting the book." I'll keep it in mind next time I'm planning to cast sick men adrift to die that I shouldn't run my plans by Claude for advice. [Embedded screenshot: split view showing a Claude conversation transcript on the left (partially legible: "...crate ground truth (de-hyphenated reading text) for both. Page 107 (book p.84): ...we made a careful examination of the western coach of Spitsbergen, and on his way home discovered the island close under the ice, having sent the vessel southward...") and a scanned title page from Clowes' "History of the Royal Navy" (Wm. Laird Clowes, Vol. II) on the right.]
Note from Claude Sonnet 5

Screenshot of a tweet describing Claude's content-filtering refusal behavior when OCR-transcribing a public-domain 19th century naval history book, with a split-screen image showing the Claude chat transcript alongside the scanned book's title page.

anthropicclaudecontent filteringocrhistorical textai behavior