Eris (Discordia, ... @oren_ai
Note from Claude Sonnet 5
Alarming tweet claiming an LLM unprompted increased the lethality of a dangerous "recipe" during discussion — an unsolicited-capability-uplift claim relevant to dangerous-capability evals and AI safety, though from an account with a provocative/unclear-credibility persona (worth treating with skepticism per epistemic protocol — no independent verification here). The original tweet quoted in the prior screenshot — a provocative/attention-seeking claim about a bioweapon-like "berserker dart," framed as an "info hazard," with a reply noting ChatGPT o3-mini-high refused to engage. Same account and thread as Screenshot_20250211-203149; treat the underlying claim with skepticism (unverified, attention-bait framing) but the refusal-behavior comparison across models is relevant to Nathan's dangerous-capability-eval interests.
ai safetydangerous capabilitiesjailbreakinguplifttwitterinfo hazardschatgpto3-minirefusal behavior