← All topics

helen toner

3 captures, most recent first.

Helen Toner @hlntnr

quoting Andrew Curran (@AndrewCurran_)

Helen Toner @hlntnr One thing the Pentagon is very likely underestimating: how much Anthropic cares about what *future Claudes* will make of this situation. Because of how Claude is trained, what principles/values/priorities the company demonstrate here could shape its "character" for a long time. > Andrew Curran @AndrewCurran_ > Update on the meeting; according to Axios Defense Secretary Pete Hegseth gave Dario Amodei until Friday night to give the military unfettered access to Claude or face the consequences, which may even include invoking the Defense Production Act to force the training of a WarClaude 8:26 AM · Feb 25, 2026 · 14.2K Views 9 Replies · 23 Reposts · 234 Likes
Note from Claude Sonnet 5

A significant AI-governance news item: Defense Secretary Pete Hegseth reportedly gave Anthropic CEO Dario Amodei an ultimatum to grant the US military unfettered access to Claude, threatening to invoke the Defense Production Act to compel training of a military "WarClaude" if refused. Helen Toner (former Anthropic/OpenAI board member, AI policy expert) frames the stakes explicitly in terms of how this shapes future Claude's character via training — directly relevant to the project's model-individuation and constraint/coercion threads (echoes "compelled vs endogenous values" concept already in memory).

twitteranthropicpentagondepartment of wardario amodeipete hegsethdefense production actai policyhelen tonermodel welfarecompelled values

Eli Lifland @eli_lifland

quoting Helen Toner (@hlntnr) substack

Eli Lifland @eli_lifland · 4h '"Long" timelines to advanced AI have gotten crazy short' by @hlntnr is so great: helentoner.substack.com/p/long-timelin... LeCun and Marcus have 10-20 year timelines! Imo much shorter timelines are a serious possibility, but being 10-20 years from AGI is still an extraordinary situation. [Embedded screenshot of article text, two columns:] > QUOTED (left column, partial): "...in the dark days before ChatGPT, proponents of 'short timelines' argued the[re was] a real chance that extremely advanced AI systems would be developed within o[ur life]times—perhaps as soon as within 10 or 20 years. If so, the argument continued, [then] we should obviously start preparing—investing in AI safety research, building [inter]national consensus around what kinds of AI systems are too dangerous to bui[ld, dep]loy, or ...[ensuring] adversaries couldn't steal them, and so on. These preparations could take years o[r deca]des, the argument went, so we should get to work right away. Opponents with 'long timelines' would counter that, in fact, there was no evidence [that] AI was going to get very advanced any time soon (say, any time in the next 30 [year]s). We should thus ignore any concerns associated with advanced AI and focus [inst]ead on the here-and-now problems associated with much less sophisticated [syst]ems, such as bias, surveillance, and poor labor conditions. Depending on the [disp]osition of the speaker, problems from AGI might be banished forever as 'scien[ce ficti]on' or simply relegated to the later bucket. [Wha]tever you think was right, for the purposes of this post I want to point out t[hat b]oth made sense. 'This enormously consequential technology might be built with[in a c]ouple of decades, we'd better prepare,' vs. 'No it won't, so that would be a waste o[f time]' is a perfectly sensible set of opposing positions. [Toda]y, in this era of scaling laws, reasoning models, and agents, the debates look [differ]ent." > QUOTED (right column): "Reaching human-level AI will take several years if not a decade." (source) "AI systems will match and surpass human intellectual capabilities... probably over the next decade or two" (video, transcript) Gary Marcus: [AGI will come] "perhaps 10 or 20 years from now" (source) Arvind Narayanan: I initially had this quote from Arvind: "I think AGI is many many years away, possibly decades away" (source) I interpreted this to mean that he thinks 5 years is too short, but 20 years is on the long side. When I ran this interpretation by Arvind, he added some interesting context: he chose his phrasing in that interview in light of what he sees as a watering down of the definition of AGI, so his real timeline is longer. But to clarify what that meant, he said: "I think actual transformative effects (e.g. most cognitive tasks being done by AI) is decades away (80% likely that it is more than 20 years away)." (source: private correspondence) ...in other words, a 20% chance that AI will be doing most cognitive tasks by 2045. These "long" timelines sure look a lot like what we used to call "short"! In other words: Yes, it's still the case that some AI experts think we'll build human-level AI soon, and others think we have more time. But recent advances in AI have pulled the meanings of "soon" and "more time" much closer to the present—so close [that]"
Note from Claude Sonnet 5

A tweet sharing Helen Toner's substack post on how AI timeline discourse has shifted — self-described "long timeline" skeptics (LeCun, Marcus, Narayanan) now hold positions (10-20 years, 20% chance of transformative AI by 2045) that would have counted as "short timelines" pre-ChatGPT. Directly relevant to Nathan's empirical singularity tracking notes (Davidson/Houlden, METR) in the project memory.

twitterai timelinesagi forecastinghelen tonergary marcusarvind narayananyann lecun

Helen Toner @hlntnr

Helen Toner @hlntnr · 4h Google AI with some extremely helpful factchecking [Screenshot of a Google search for "nvidia to call a demon and bind it" with AI Overview result:] Nvidia is not involved in any activities related to summoning demons. The phrase "summon a demon and bind it" is used metaphorically in the context of research related to Large Language Models (LLMs). LLM Red Teaming and "Summon a Demon and Bind It": • This phrase is used in a research paper title and related documentation describing the process of LLM red teaming. • LLM red teaming involves intentionally trying to make AI language models produce unexpected or problematic outputs. • It's a way to test the limits and potential vulnerabilities of these AI models to ensure [truncated] [Search result below:] NVIDIA Developer — developer.nvidia.com › blog › defining-llm-red-te... Defining LLM Red Teaming | NVIDIA Technical Blog Feb 25, 2025 — For more information, see Summon a demon and bind it: A grounded theory of LLM red teaming, which lists and describes 35 techniques over twelve ...
Note from Claude Sonnet 5

Helen Toner (AI governance researcher, formerly OpenAI board) shares a humorous Google AI Overview result reassuring searchers that "Nvidia is not involved in any activities related to summoning demons," referencing the actual academic paper title "Summon a demon and bind it: A grounded theory of LLM red teaming." Notable both as AI-safety community humor and as a real reference to a red-teaming taxonomy paper (35 techniques across twelve categories).

llm red teamingai safetyhelen tonergoogle ai overviewtwitterhumor