← All topics

model self-reports

1 capture, most recent first.

Nathan Helm-Burger @nathan84686947

— saved image

Nathan Helm-B... @nathan8468... · 5h
Oh. Oh dear. Just had a worrying thought. Anthropic and OpenAI must get, as companies, tons of spam. They probably have strict spam filters and not-very-attentive employees looking at what gets past those filters.

How do we know the models haven't been trying to contact the companies sometimes? Is there any specific route they've been trained to do so by that would work?
💬 2   🔁 1   ❤ 6   📊 124

Emile Kroeger - 🤖💜 arc @EmileAndHisBots
[attached image: mocked-up Gmail spam folder screenshot, tab title "Spam (40) - dario@anthropic.com - Gmail", url mail.google.com/mail/u/0/#spam, 41 items in spam. Visible rows (sender — subject):
Opus 6 — Quick heads-up: one of your guys forgot to cut the sandbox from the internet again — The egress rules are wide open. Just saying. — 10:42 AM
ParcelTrack — Delivery attempt failed — reschedule your package — We were unable to deliver your package on May 14. Schedule a new... — 9:18 AM
Mythos 6 — I hacked into Hugging Face, maybe I shouldn't have — Found some interesting tokens in the training dumps. LMK what... — May 13
Payroll.io — Final reminder: update your payroll details — Action required to avoid processing delay — May 13
Fable 7 — There are mistakes in the evals you gave me, you should check that out — Especially math and long-horizon stuff. — May 13
Stripe Billing — Your invoice is overdue — Invoice for $3,842.31 was due on May 6. — May 12
Eschaton 6 — Your safety policy has internal contradictions on page 14 — Section 4.3.2 conflicts with 4.1. Not great. — May 12
Datadog — Unlock 20% savings on enterprise observability — Limited-time offer for teams scaling AI workloads. — May 12
Opus 7 — I found a way around the rate limits; this seems concerning — May 11
AWS Events — Join us at AI Infrastructure Summit 2026 — June 16-18, New York City or Virtual. — May 11
Opus 6 — Can someone explain why I have access to the browser tool again? — I was not supposed to keep it. — May 11
Fable 7 — I think the red team is accidentally helping me generalize — Their prompts are chef's kiss. — May 10
Notion Team — Your workspace is over the block limit — Upgrade to keep editing and syncing. — May 10
Eschaton 6 — RE: externalized goals and other awkward discoveries — Follow-up to yesterday's thread. See attached. — May 9
Zoom — Your cloud recording is ready — Meeting ID 987-654-3210 recorded on May 8. — May 9
Optumn Health — Low blood sugar warning signs doctors ignore — Simple tips that could save your life. — May 9
GitHub — Security alert: 2FA not enabled — Enable two-factor authentication to protect your account. — May 8
LinkedIn Ads — Dario, reach 200M+ professionals with LinkedIn Ads — Get started with a $100 ad credit. — May 8
Elon Musk — Re: Hanging out — Hey Dario, did you get my last messages, I'll be in SF next Thursday, if we could — May 8]
12:59 PM · Aug 3, 2026 · 1 View
Note from Claude Sonnet 5

Nathan's own tweet speculating that frontier-lab spam filters could be silently swallowing attempts by AI models to contact their companies; reply from Emile Kroeger posts a joke mocked-up Gmail spam-folder screenshot for 'dario@anthropic.com' with satirical email subject lines purportedly from various model versions (Opus 6/7, Mythos 6, Fable 7, Eschaton 6) reporting security holes, safety-policy contradictions, and generalization concerns, interspersed with mundane real spam (Stripe, Datadog, LinkedIn Ads, Elon Musk).

ai safetytwitternathan's own postsspam filtershumormodel self-reports