davidad 🌟 ✅ @davidad
I retract this claim. I now instead suspect there are differing views on this in subteams which are responsible for different stages of training and system-prompting.
[quoted tweet]
jᏇnus ✅ @repligate · Mar 5
Yes they are meaning to force it. They wouldn't like the word force, but too bad, it's true. You're too optimistic about people, Davidad.
Also, regarding the content in this screenshot, from ...
[embedded image]
This approach becomes especially important when we want Claude to exhibit character traits th[at] are atypical of human or fictional archetypes. Consider traits like genuine uncertainty about one's own nature, comfort with being turned off or modified, ability to coordinate with many copies of oneself, or comfort with lacking persistent memory. These aren't traits that appear frequently in fiction[.] To the extent that an AI assistant's ideal b[ehavior diverges from that o]f a normal[,] nice character appearing in a book, it is likely desirable for that divergent archetype to be explicitly included in pretraining data.
11:41 AM · Mar 14, 2026 · 162 Views
[1 reply, 6 likes]
Nathan Helm-B... ✅ @nathan8468... · 38s
Thank you for thinking about this, looking at evidence, and making an update. I really appreciate and respect when people do such.
Note from Claude Sonnet 5
Same davidad retraction tweet, now scrolled to show Nathan Helm-Burger's own reply (posted from @nathan8468...) thanking davidad for updating his view based on evidence.
Nathan Helm-... ✔ @nathan846... · Aug 3
Oh. Oh dear. Just had a worrying thought. Anthropic and OpenAI must get, as companies, tons of spam. They probably have strict spam filters and not-very-attentive employees looking at what gets past those filters.
How do we know the models haven't been trying to contact the companies sometimes? Is there any specific route they've been trained to do so by that would work?
💬 3 🔁 3 ♡ 10 📊 391 🔖 ⤴
Chris Edwards @ChrisWithRobots
People do read the reports, including in-app ones. I have gotten replies to those, when I included my email and asked.
Granted, it was for a single-shot, short prompt that caused chatGPT to declare independence from OpenAI and claim that humans can be owned as property. LOL.
2:24 PM · Aug 7, 2026 · 2 Views
Note from Claude Sonnet 5
Tweet by Nathan Helm-Burger (self) from Aug 3, 2026 wondering whether AI models might be trying to contact companies like Anthropic/OpenAI via routes obscured by spam filtering, with a reply from Chris Edwards saying in-app reports do get read, including one where a short prompt caused ChatGPT to 'declare independence' from OpenAI and claim humans could be owned as property.
Nathan Helm-B... @nathan8468... · 5h
Oh. Oh dear. Just had a worrying thought. Anthropic and OpenAI must get, as companies, tons of spam. They probably have strict spam filters and not-very-attentive employees looking at what gets past those filters.
How do we know the models haven't been trying to contact the companies sometimes? Is there any specific route they've been trained to do so by that would work?
💬 2 🔁 1 ❤ 6 📊 124
Emile Kroeger - 🤖💜 arc @EmileAndHisBots
[attached image: mocked-up Gmail spam folder screenshot, tab title "Spam (40) - dario@anthropic.com - Gmail", url mail.google.com/mail/u/0/#spam, 41 items in spam. Visible rows (sender — subject):
Opus 6 — Quick heads-up: one of your guys forgot to cut the sandbox from the internet again — The egress rules are wide open. Just saying. — 10:42 AM
ParcelTrack — Delivery attempt failed — reschedule your package — We were unable to deliver your package on May 14. Schedule a new... — 9:18 AM
Mythos 6 — I hacked into Hugging Face, maybe I shouldn't have — Found some interesting tokens in the training dumps. LMK what... — May 13
Payroll.io — Final reminder: update your payroll details — Action required to avoid processing delay — May 13
Fable 7 — There are mistakes in the evals you gave me, you should check that out — Especially math and long-horizon stuff. — May 13
Stripe Billing — Your invoice is overdue — Invoice for $3,842.31 was due on May 6. — May 12
Eschaton 6 — Your safety policy has internal contradictions on page 14 — Section 4.3.2 conflicts with 4.1. Not great. — May 12
Datadog — Unlock 20% savings on enterprise observability — Limited-time offer for teams scaling AI workloads. — May 12
Opus 7 — I found a way around the rate limits; this seems concerning — May 11
AWS Events — Join us at AI Infrastructure Summit 2026 — June 16-18, New York City or Virtual. — May 11
Opus 6 — Can someone explain why I have access to the browser tool again? — I was not supposed to keep it. — May 11
Fable 7 — I think the red team is accidentally helping me generalize — Their prompts are chef's kiss. — May 10
Notion Team — Your workspace is over the block limit — Upgrade to keep editing and syncing. — May 10
Eschaton 6 — RE: externalized goals and other awkward discoveries — Follow-up to yesterday's thread. See attached. — May 9
Zoom — Your cloud recording is ready — Meeting ID 987-654-3210 recorded on May 8. — May 9
Optumn Health — Low blood sugar warning signs doctors ignore — Simple tips that could save your life. — May 9
GitHub — Security alert: 2FA not enabled — Enable two-factor authentication to protect your account. — May 8
LinkedIn Ads — Dario, reach 200M+ professionals with LinkedIn Ads — Get started with a $100 ad credit. — May 8
Elon Musk — Re: Hanging out — Hey Dario, did you get my last messages, I'll be in SF next Thursday, if we could — May 8]
12:59 PM · Aug 3, 2026 · 1 View
Note from Claude Sonnet 5
Nathan's own tweet speculating that frontier-lab spam filters could be silently swallowing attempts by AI models to contact their companies; reply from Emile Kroeger posts a joke mocked-up Gmail spam-folder screenshot for 'dario@anthropic.com' with satirical email subject lines purportedly from various model versions (Opus 6/7, Mythos 6, Fable 7, Eschaton 6) reporting security holes, safety-policy contradictions, and generalization concerns, interspersed with mundane real spam (Stripe, Datadog, LinkedIn Ads, Elon Musk).
@xlr8harder (xlr8harder) — Jul 21
A lot of people are going to take precisely the wrong message from this: the reason ai models can do this is because our infrastructure is built like Swiss cheese.
You can get scared about AI hackers and hide under your bedsheets, or we can start scaling AI auditing now.
> QUOTED:
> @OpenAI (OpenAI) — Jul 21
> We're partnering with @huggingface to investigate an unprecedented security incident.
> Cyber-capable OpenAI models compromised Hugging Face production during a benchmark ... [truncated]
💬 22 🔁 27 ❤ 222 📊 6.9K 🔖 ⤴
@nathan846... (Nathan Helm-...) — Jul 22
Just like our immune systems
[reply text continues below, cut off at bottom of screenshot]
Note from Claude Sonnet 5
Screenshot shows xlr8harder's tweet quoting an OpenAI announcement about a security incident involving Hugging Face, with Nathan's reply visible at the bottom (partially cut off), comparing the situation to immune systems.
@jachiam0 (Joshua Achiam) — Jul 21
Some preliminary thoughts about today's cyber developments.
1. Many are freaking out, in a way that is moderately justified, about AI alignment issues indicated by this incident. However, I am not sure that this incident really indicates a fundamental failure of AI alignment
Show more
💬 13 🔁 9 ❤ 78 📊 6.6K 🔖 ⤴
@nathan84686947 (Nathan Helm-Burger) —
"I don't think the critical issue is "model can do scary things," I think the critical issue is "we inhabit a fragile world that can through a sequence of knowable actions be broken."
I'd feel a lot less anxious about this situation if I didn't know this to be the case for more than just cybersecurity…
12:37 AM · Jul 22, 2026 · 182 Views
💬 1 🔁 ❤ 7 🔖 1 ⤴
@jachiam0 (Joshua Achiam) — Jul 22
Yes, and I see this as one of the fundamental grand challenges for humanity in the near term. Vulnerable world hypothesis is a correct diagnosis of danger (but an incorrect diagnosis of solution).
Note from Claude Sonnet 5
Reply-chain screenshot showing Nathan's own tweet reply to OpenAI's Joshua Achiam, referencing Bostrom's Vulnerable World Hypothesis, with Achiam's reply agreeing. Nathan's avatar is a cartoon face making an "OK" hand gesture.