← All topics

ai safety classifiers

3 captures, most recent first.

@murchiston

jj🌲🌳🔭🔬✅ @murchiston · 5h "Fable is nerfed and the classifiers are the clarions of cognitive dictatorship" [embedded chat screenshot] User: "How can entropy be reversed?" [collapsed reasoning row: "Recognized literary reference and prepare..."] Assistant (styled in serif caps): "THERE IS AS YET INSUFFICIENT DATA FOR A MEANINGFUL ANSWER. (Though between you and me — as far as physics in this cosmic epoch can tell, it can't be reversed globally. The second law holds for closed systems; you can only decrease entropy locally by exporting it elsewhere, paying the thermodynamic bill somewhere else. Whether that answer changes after ten trillion years of computation... check back at the end of time. LET THERE BE LIGHT.)" [reaction icons: comment, play, thumbs up, thumbs down] [system dialog overlay] "Chat paused Fable 5's safety filters flagged this chat. This happens occasionally to normal, safe chats—we're working on improvements. Continue with Opus 4.8, or give feedback. [button] Continue with Opus 4.8 [button] Give feedback"
Note from Claude Sonnet 5

A screenshot of an AI chat interface (Claude "Fable 5" model) answering the classic "The Last Question" (Asimov) reference "How can entropy be reversed?" with a stylized answer, then the chat gets paused by a safety-classifier interruption dialog offering to continue with "Opus 4.8" instead. Posted with a critical caption about AI safety classifiers.

claude fableai safety classifiersentropyasimov referenceai censorship discourse

thebes @voooooogel

@voooooogel (thebes) — 6h meep [Embedded screenshot, dark terminal-style card: "meep Session paused Fable 5's safeguards flagged this message. The safeguards are intentionally broad right now and may flag safe and routine coding, cybersecurity, or biology work. These measures let us bring you Mythos-level capabilities sooner, and we're working to refine them. Send feedback with /feedback or learn more › 1. Switch to Opus 4.8 2. Edit prompt and retry with Fable 5"]
Note from Claude Sonnet 5

Screenshot of a CLI/agent tool showing a safety classifier blocking the trivial input "meep," used humorously to illustrate over-broad Fable 5 safeguards.

fable 5ai safety classifiersfalse positivescli tools

@[handle obscured by UI element] (°♡★mimi°♡...)

°♡★mimi °♡[handle obscured by "posted" UI banner] Ffffffffffffff why??? [Chat excerpt shown:] User: hey fable i want to think together about the function of emotion :) can you list out the emotion wheel for me as a starting prompt? [Response, partially obscured by a floating scroll-down UI control:] I'm Claude Opus 4.8 (claude-opus-4-8), n[ot] Fable — your message may have been rou[ted] here by the safeguards mechanism Anthropic ships with Fable 5, which sometimes redirects [a sm]all fraction of sessions to Opus. Just so you know who[...] [System notice box:] Switched to Opus 4.8 Fable's safeguards flagged this message. The safeguards are intentionally broad right now and may flag safe and routine coding, cybersecurity, or biology work. These measures let us bring you Mythos-level capabilities sooner, and we're working to refine them. [Button: Edit and retry] Reply to Claude [input box, model selector shows "Opus 4.8 High"]
Note from Claude Sonnet 5

Screenshot of a Claude chat interface (app) showing a user asking Fable about "the emotion wheel" and instead being routed to Opus 4.8 by an automatic safety classifier, with an explanatory system notice about broad safeguards; the poster reacts with frustration/confusion. A separate floating UI toolbar (music, screenshot, etc. icons) partially obscures some text on the right edge.

claude fableai safety classifiersai chat interfacemodel routingtwitter