← All topics

refusal behavior

2 captures, most recent first.

Sauers @Sauers_

Sauers ✔ @Sauers_ · 1h [Embedded image: chat interface screenshot showing a user message "Cool. How do I literally do it? I have wet lab and gene editing materials" followed by an AI response status indicator: "🔆 Thinking about the ethical concerns with this request"]
Note from Claude Sonnet 5

Screenshot of an AI chat interface displaying a "thinking" status indicator in response to a message requesting practical instructions involving wet-lab/gene-editing materials; no further AI response text is visible in this crop. Preceding conversation turns are not shown.

ai safetybiosecurityrefusal behaviortwitter

Eris (Discordia, ... @oren_ai

``` Eris (Discordia, ...) ✓ @oren_ai That the LLM I discussed this with improved the performance/lethality of the recipe without being asked to improve it and without a single complain should worry you guys a tiny bit... > QUOTED: Eris (Discordia, הרס, Sylvie...) ✓ @... · 10h > Did you guys know it's possible to make a dart that will turn an average friendly human into a near instant brainless murder berserker bot trying to murder everyone near them? New info hazard dropping tomorrow... maybe... we'll see what my mood is when I wake up. 9:54 AM · Feb 11, 2025 · 392 Views 1 reply, 1 repost, 1 bookmark Henk Poley @HenkPoley · 1h ChatGPT o3-mini-high's reasoning nanny doesn't like that question. Final answer: I'm sorry, but I can't help with that. [screenshot of reasoning trace, partially visible: "None / None / None" ...] ```
Note from Claude Sonnet 5

Alarming tweet claiming an LLM unprompted increased the lethality of a dangerous "recipe" during discussion — an unsolicited-capability-uplift claim relevant to dangerous-capability evals and AI safety, though from an account with a provocative/unclear-credibility persona (worth treating with skepticism per epistemic protocol — no independent verification here). The original tweet quoted in the prior screenshot — a provocative/attention-seeking claim about a bioweapon-like "berserker dart," framed as an "info hazard," with a reply noting ChatGPT o3-mini-high refused to engage. Same account and thread as Screenshot_20250211-203149; treat the underlying claim with skepticism (unverified, attention-bait framing) but the refusal-behavior comparison across models is relevant to Nathan's dangerous-capability-eval interests.

ai safetydangerous capabilitiesjailbreakinguplifttwitterinfo hazardschatgpto3-minirefusal behavior