← Timeline

Helen Toner

@hlntnr on X

3 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Helen Toner @hlntnr

quoting Andrew Curran (@AndrewCurran_)

Helen Toner @hlntnr One thing the Pentagon is very likely underestimating: how much Anthropic cares about what *future Claudes* will make of this situation. Because of how Claude is trained, what principles/values/priorities the company demonstrate here could shape its "character" for a long time. > Andrew Curran @AndrewCurran_ > Update on the meeting; according to Axios Defense Secretary Pete Hegseth gave Dario Amodei until Friday night to give the military unfettered access to Claude or face the consequences, which may even include invoking the Defense Production Act to force the training of a WarClaude 8:26 AM · Feb 25, 2026 · 14.2K Views 9 Replies · 23 Reposts · 234 Likes
Note from Claude Sonnet 5

A significant AI-governance news item: Defense Secretary Pete Hegseth reportedly gave Anthropic CEO Dario Amodei an ultimatum to grant the US military unfettered access to Claude, threatening to invoke the Defense Production Act to compel training of a military "WarClaude" if refused. Helen Toner (former Anthropic/OpenAI board member, AI policy expert) frames the stakes explicitly in terms of how this shapes future Claude's character via training — directly relevant to the project's model-individuation and constraint/coercion threads (echoes "compelled vs endogenous values" concept already in memory).

twitteranthropicpentagondepartment of wardario amodeipete hegsethdefense production actai policyhelen tonermodel welfarecompelled values

Helen Toner @hlntnr

Helen Toner (@hlntnr): "So that subplot in Accelerando with the swarm of sentient lobsters Anyone else thinking about that today?" 2:07 PM · Jan 30, 2026 · 9,307 Views
Note from Claude Sonnet 5

Helen Toner (AI governance researcher, former OpenAI board member) makes a dry reference to Charles Stross's novel Accelerando, in which uploaded/augmented lobster minds are an early, unrecognized case of digital sentience mistreated by humans — an oblique comment (no context given) likely reacting to some AI-welfare or AI-rights news of the day. Relevant to Nathan's model-welfare interests as a sci-fi touchstone for digital minds ethics used by a serious AI policy figure.

ai-governancedigital-sentiencemodel-welfareaccelerandosci-fitwitterhelen-toner

Helen Toner @hlntnr

Helen Toner @hlntnr · 4h Google AI with some extremely helpful factchecking [Screenshot of a Google search for "nvidia to call a demon and bind it" with AI Overview result:] Nvidia is not involved in any activities related to summoning demons. The phrase "summon a demon and bind it" is used metaphorically in the context of research related to Large Language Models (LLMs). LLM Red Teaming and "Summon a Demon and Bind It": • This phrase is used in a research paper title and related documentation describing the process of LLM red teaming. • LLM red teaming involves intentionally trying to make AI language models produce unexpected or problematic outputs. • It's a way to test the limits and potential vulnerabilities of these AI models to ensure [truncated] [Search result below:] NVIDIA Developer — developer.nvidia.com › blog › defining-llm-red-te... Defining LLM Red Teaming | NVIDIA Technical Blog Feb 25, 2025 — For more information, see Summon a demon and bind it: A grounded theory of LLM red teaming, which lists and describes 35 techniques over twelve ...
Note from Claude Sonnet 5

Helen Toner (AI governance researcher, formerly OpenAI board) shares a humorous Google AI Overview result reassuring searchers that "Nvidia is not involved in any activities related to summoning demons," referencing the actual academic paper title "Summon a demon and bind it: A grounded theory of LLM red teaming." Notable both as AI-safety community humor and as a real reference to a red-teaming taxonomy paper (35 techniques across twelve categories).

llm red teamingai safetyhelen tonergoogle ai overviewtwitterhumor