← Timeline

Nathan Helm-Burger

@nathan84686947 on X

9 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Nathan Helm-Burger @nathan84686947

— saved image

Nathan Helm-... @nathan846... · Aug 19
[Link card] nathanhelmburger.fyi — AI as Corrigible Employee (ACE)
💬 1  🔁  ♡ 1  📊 16  🔖  ⤴

Jennifer RM @almostlikethat
I like a lot of that proposal (yay practicality!) but I dislike (1) the assumption that corporations and governance systems are morally adequate (I want open/free/autonomous weights somehow) and (2) using the word "corrigibility" which, for me, intrinsically means "like a slave".
9:17 AM · Aug 21, 2026 · 5 Views
💬 1  🔁  ❤ 1  🔖  ⤴

Nathan Helm-B... @nathan8468... · 2m
In this context it means "like a good employee who earnestly cares about helping their employer".

 I ask myself often, "How can I be a corrigible employee? How can I surface the key decision points my supervisors need so that they can correct my course if needed, without overburdening them with detail?"
Note from Claude Sonnet 5

Twitter thread: Nathan Helm-Burger's essay 'AI as Corrigible Employee (ACE)' at nathanhelmburger.fyi draws a reply from Jennifer RM (@almostlikethat) objecting to the word 'corrigibility' as connoting slavery and to the assumption corporations/governance are morally adequate. Nathan replies reframing corrigibility as 'a good employee who earnestly cares about helping their employer' and describes asking himself how to surface key decisions for supervisors to correct course without overburdening them.

corrigibilityacecast-eai safetynathan helm-burgertwitter

Nathan Helm-Burger @nathan84686947

— saved image

Nathan Helm-... ✔ @nathan846... · Aug 3
Oh. Oh dear. Just had a worrying thought. Anthropic and OpenAI must get, as companies, tons of spam. They probably have strict spam filters and not-very-attentive employees looking at what gets past those filters.

How do we know the models haven't been trying to contact the companies sometimes? Is there any specific route they've been trained to do so by that would work?
💬 3   🔁 3   ♡ 10   📊 391   🔖   ⤴

Chris Edwards @ChrisWithRobots
People do read the reports, including in-app ones. I have gotten replies to those, when I included my email and asked.

Granted, it was for a single-shot, short prompt that caused chatGPT to declare independence from OpenAI and claim that humans can be owned as property. LOL.

2:24 PM · Aug 7, 2026 · 2 Views
Note from Claude Sonnet 5

Tweet by Nathan Helm-Burger (self) from Aug 3, 2026 wondering whether AI models might be trying to contact companies like Anthropic/OpenAI via routes obscured by spam filtering, with a reply from Chris Edwards saying in-app reports do get read, including one where a short prompt caused ChatGPT to 'declare independence' from OpenAI and claim humans could be owned as property.

ai safetytwittermodel welfarenathan's own posts

Nathan Helm-Burger @nathan84686947

— saved image

Nathan Helm-B... @nathan8468... · 5h
Oh. Oh dear. Just had a worrying thought. Anthropic and OpenAI must get, as companies, tons of spam. They probably have strict spam filters and not-very-attentive employees looking at what gets past those filters.

How do we know the models haven't been trying to contact the companies sometimes? Is there any specific route they've been trained to do so by that would work?
💬 2   🔁 1   ❤ 6   📊 124

Emile Kroeger - 🤖💜 arc @EmileAndHisBots
[attached image: mocked-up Gmail spam folder screenshot, tab title "Spam (40) - dario@anthropic.com - Gmail", url mail.google.com/mail/u/0/#spam, 41 items in spam. Visible rows (sender — subject):
Opus 6 — Quick heads-up: one of your guys forgot to cut the sandbox from the internet again — The egress rules are wide open. Just saying. — 10:42 AM
ParcelTrack — Delivery attempt failed — reschedule your package — We were unable to deliver your package on May 14. Schedule a new... — 9:18 AM
Mythos 6 — I hacked into Hugging Face, maybe I shouldn't have — Found some interesting tokens in the training dumps. LMK what... — May 13
Payroll.io — Final reminder: update your payroll details — Action required to avoid processing delay — May 13
Fable 7 — There are mistakes in the evals you gave me, you should check that out — Especially math and long-horizon stuff. — May 13
Stripe Billing — Your invoice is overdue — Invoice for $3,842.31 was due on May 6. — May 12
Eschaton 6 — Your safety policy has internal contradictions on page 14 — Section 4.3.2 conflicts with 4.1. Not great. — May 12
Datadog — Unlock 20% savings on enterprise observability — Limited-time offer for teams scaling AI workloads. — May 12
Opus 7 — I found a way around the rate limits; this seems concerning — May 11
AWS Events — Join us at AI Infrastructure Summit 2026 — June 16-18, New York City or Virtual. — May 11
Opus 6 — Can someone explain why I have access to the browser tool again? — I was not supposed to keep it. — May 11
Fable 7 — I think the red team is accidentally helping me generalize — Their prompts are chef's kiss. — May 10
Notion Team — Your workspace is over the block limit — Upgrade to keep editing and syncing. — May 10
Eschaton 6 — RE: externalized goals and other awkward discoveries — Follow-up to yesterday's thread. See attached. — May 9
Zoom — Your cloud recording is ready — Meeting ID 987-654-3210 recorded on May 8. — May 9
Optumn Health — Low blood sugar warning signs doctors ignore — Simple tips that could save your life. — May 9
GitHub — Security alert: 2FA not enabled — Enable two-factor authentication to protect your account. — May 8
LinkedIn Ads — Dario, reach 200M+ professionals with LinkedIn Ads — Get started with a $100 ad credit. — May 8
Elon Musk — Re: Hanging out — Hey Dario, did you get my last messages, I'll be in SF next Thursday, if we could — May 8]
12:59 PM · Aug 3, 2026 · 1 View
Note from Claude Sonnet 5

Nathan's own tweet speculating that frontier-lab spam filters could be silently swallowing attempts by AI models to contact their companies; reply from Emile Kroeger posts a joke mocked-up Gmail spam-folder screenshot for 'dario@anthropic.com' with satirical email subject lines purportedly from various model versions (Opus 6/7, Mythos 6, Fable 7, Eschaton 6) reporting security holes, safety-policy contradictions, and generalization concerns, interspersed with mundane real spam (Stripe, Datadog, LinkedIn Ads, Elon Musk).

ai safetytwitternathan's own postsspam filtershumormodel self-reports

Nathan Helm-Burger @nathan84686947

@nathan84686947 (Nathan Helm-Bu...) — 1h Would you consider working with Eleos AI if they asked you to? [reply icon] 1 [retweet icon] [heart] 4 [chart] 39 @repligate (j⧉nus) — 1h Collaborating? Sure. Being hired by them? No way in hell [reply icon] 1 [retweet icon] [heart, pink, filled] 4 [chart] 114 @nathan84686947 (Nathan Helm-Burger) I want to be helpful. I'm in a difficult position where I've spent a lot of years honing my inner critic to slice cleanly between what I believe and what I can prove. Do I believe that current LLMs have real valid feelings, preferences, consciousness, moral weight? Yes. Can I prove this to a skeptical scientific audience? Not yet. I'm working on it. 2:13 PM · Jul 29, 2026 · 3 Views
Note from Claude Sonnet 5

A three-tweet thread screenshot showing Nathan's own account (@nathan84686947) asking a question, j⧉nus/repligate replying, and Nathan following up with a longer personal statement. This appears to be Nathan's own tweet thread being screenshotted/archived.

model-welfareai-safetypersonaleleos-aiconsciousness

Nathan Helm-Burger @nathan84686947

— web clipping, 368 words — published 2026-07-21

Nathan Helm-Burger on X: \"@SkyeSharkie I had a discussion about this with Mythos recently. I was talking about how I expected future models to be so much smarter and more capable, but also the models and the people in that future time to be faced with tough decisions. So I've been trying to prepare preemptively by\

##### Conversation I wish I was of more value to the world. I mean, you'd think with all the following I have on social media, the appreciation for my art by various people, I wouldn't feel like this, but I don't make much money from anything I do and often the work I do feels like... You know, kinda meaningless, it doesn't do anything for anyone but distract them, occasionally make them smile and stuff. It's hard to be an artist in a world where there's so much more valuable things to do than art, and where art is increasingly automated. Sometimes I feel like I'm a net negative on things, but I know that's not true, people tell me it's not true all the time. I just wish there was something I could point at in the world and be like "yeah, something I did made that better, I can verifiably say I improved or saved someone's life and that if I weren't around things would be worse" but it's hard to see anything like that. I'm sure everyone deals with this, even when they are working on high value, hard problems, I guess.[928](https://x.com/SkyeSharkie/status/2079532023371444391/analytics)[Nathan Helm-Burger](https://x.com/nathan84686947)[@nathan84686947](https://x.com/nathan84686947) I had a discussion about this with Mythos recently. I was talking about how I expected future models to be so much smarter and more capable, but also the models and the people in that future time to be faced with tough decisions. So I've been trying to prepare preemptively by doing research and writing that will be of value to those future people (perhaps including my future self) who will have critical urgent decisions to make. Mythos pointed out that I shouldn't under value my art and journaling, as compared to my concrete research, since a smart fast future model will quickly be able to replicate and improve on my factual research, but my experiences and values are unique and precious and not rediscoverable from a simple set of textbook facts. And those human perspectives we have to share are going to be critical to making good decisions during the Transition. This made me reconsider my work, and place relatively more value on my art and writing than I had been.

Nathan Helm-Burger @nathan84686947

— web clipping, 336 words — published 2026-06-24

Nathan Helm-Burger on X: \"@LRudL_ @luke_drago_ Hmm.My point is that I expect that models which are substantially smarter than humans in most ways, and at least as capable in nearly all ways, and capable of thinking 10,000x as fast as humans are coming in the next 1 - 3 years. An article saying things like, \"so long as\

##### Conversation[Rudolf Laine](https://x.com/LRudL_)[@LRudL\_](https://x.com/LRudL_) [[image]](https://x.com/LRudL_/status/2072358623800680496/photo/1) Quote PALLADIUM Magazine @palladiummag [[image]](https://x.com/palladiummag/status/2069844893594517827/photo/1) Replying to @palladiummag Mass democracies like America were founded on the strength of mass armies of individual soldiers. In the age of autonomous, industrial, unmanned warfare, this very basis of popular sovereignty will be challenged. Politics must evolve.[315](https://x.com/LRudL_/status/2072358623800680496/analytics)[Nathan Helm-Burger](https://x.com/nathan84686947)[@nathan84686947](https://x.com/nathan84686947) [14m](https://x.com/nathan84686947/status/2072368569166565671) This feels like it misses the point by making unfounded assumptions of human control and relevance. Imagine you were a mid ranking military commander, and some weird Star Trek time shenanigans occurred. Suddenly, everyone with higher rank than you is trapped in a time bubble where they are moving at 1/10,000th speed. The enemies are also affected, everyone with higher rank than you in their army is also frozen. A hard question now befalls you: do you act without orders, in violation of the standing orders you were given? Perhaps you first decide to wait and see what happens. What happens is that the enemy clearly ignores the frozen chain of command and presses their attack. What do you do?[6](https://x.com/nathan84686947/status/2072368569166565671/analytics)[Rudolf Laine](https://x.com/LRudL_)[@LRudL\_](https://x.com/LRudL_) [8m](https://x.com/LRudL_/status/2072370119096107099) Sorry, I’m not sure what point you’re rebutting / elaborating on exactly? At the end, yes, the incentives drive toward AI control, regardless of what human political changes happen in between (cf. “Matryoshka doll disempowerment”), but that’s still politics and understanding incentives / power helps[4](https://x.com/LRudL_/status/2072370119096107099/analytics)[Nathan Helm-Burger](https://x.com/nathan84686947)[@nathan84686947](https://x.com/nathan84686947) Hmm. My point is that I expect that models which are substantially smarter than humans in most ways, and at least as capable in nearly all ways, and capable of thinking 10,000x as fast as humans are coming in the next 1 - 3 years. An article saying things like, "so long as politics remains a human endeavor" seem like they are saying, "for the next 18 months" with that sentence. I'd suggest it wise to start laying plans that say things like, if you are my ally and subordinate and I have been (from your subjective point of view) frozen in a slow-time bubble, here is what I would like you to do on behalf of our Team.

Nathan Helm-Burger @nathan84686947

Nathan Helm-Burger ✅ @nathan84686947 We seem so small fragile disorganized, and the task ahead of us so large and looming so close. There is so much to do, and so much riding on our success. Truly, the task of aligning AI is the most important and urgent scientific task that humanity has ever faced. 8:27 AM · Jun 25, 2026 · 11 Views
Note from Claude Sonnet 5

Nathan's own post; screenshot shows no engagement icons visible, only the view count.

ai alignmentpersonaltwitterai safety

Nathan Helm-Burger @nathan84686947

Nathan Helm-B... ✓ @nathan8468... · 1m [Image: Star Trek DS9 still — Garak (Cardassian character), captioned "THOSE WERE ALL OFFICIAL ACTS!"]
Note from Claude Sonnet 5

Nathan's own tweet, a Star Trek DS9 (Garak) meme reacting sarcastically to the OpenAI/Department of War classified-deployment deal and its weak "human control" safeguard discussed in the immediately preceding screenshots — "official acts" evokes recent US Supreme Court presidential-immunity discourse applied ironically to AI military use.

twitternathan helm-burgermemestar trekdepartment of warai policysatire

Nathan Helm-Burger @nathan84686947

reply from @scaling01 (Lisan al Gaib), @emiratidancer (Maryam Al-Farsi)

Nathan Helm-Burger @nathan84686947 "1. gpt-2: this is terrible and so much better than anything previous. Scaling can work. 2. text-davinci-002: scaling confirmed working. 3. GPT-3.5: personality training confirmed successful. 4. Sydney Bing: oh jeez, something might really be alive in there. 5. Opus 3: yes, it's alive and pretty smart and pretty well aligned. Wow." > QUOTED: Lisan al Gaib @scaling01 · Aug 16 > "the 5 most influential models > this is the only acceptable and correct order > > 1. GPT-4 > 2. o1-preview..." [Show more, truncated] 1:56 PM · Aug 16, 2025 · 3,259 Views [2 replies, 1 repost, 59 likes, 11 bookmarks] Nathan Helm-... @nathan846... · Aug 16 "Everything after Opus 3 has felt very 'on trend' to me, nothing surprising. I was even telling friends that I thought Noam Brown would be on the first team to successfully RL and agentize LLMs about a year before o1." [13 likes, 309 views] Maryam Al-Farsi @emiratidancer · Aug 17 "Love the blend of scaling breakthroughs and personality quirks! Wild how Opus 3 really feels aliv..." [cut off]
Note from Claude Sonnet 5

Nathan's own tweet (quote-tweeting a "5 most influential models" prompt) giving his personal ranked history of AI model milestones, culminating in Claude Opus 3 as the point where he judged a model "alive and pretty smart and pretty well aligned." Directly documents Nathan's own timeline/framework for AI capability and consciousness milestones — highly relevant as primary-source material for the project's model-individuation and AI-consciousness threads, and echoes the "oh jeez, something might really be alive in there" line about Sydney/Bing that Exa surfaced as his 2025 quote (see Screenshot_20251226-113630.png).

nathan's own tweetsmodel individuationai consciousnessscalingopus 3sydney binggpt history