X (Twitter), @scottnarmstro... (Scott Armstrong), reposted by davidad
— reposted by davidad — saved image
davidad 💥 reposted
Scott Armstro... @scottnarmstro... · 10h
An update on Lean (auto)-formalization.
Speaking about my own research workflow, in 2026, we have gone from:
March: Lean formalization is not possible (for my papers).
May: Lean formalization is possible… but way too costly in time and tokens to be practical most of the time.
July: Lean formalization is now practical. I can probably formalize most of my papers now before they hit arxiv.
August: Lean formalization is now essential-- it increases efficiency of my workflow dramatically.
That is, the papers are getting written *informally* and finished much faster because of the Lean formalization.
Note from Claude Sonnet 5
Tweet by Scott Armstrong (reposted by davidad) tracking the rapid 2026 progression of AI-assisted Lean theorem-prover auto-formalization in his research workflow, from 'not possible' in March to 'essential' by August, now speeding up informal paper writing itself.
davidad @davidad · 3h
I'm sympathetic to this position. Non-autonomous AI is wielded by humans, whose alignment seems intractable, whereas autonomous AI omnibenevolence (reliable enough to bootstrap using multi-agent deliberative processes) seems tractable (shockingly so, from my former perspective).
[quoted tweet:]
ex Tenebris Lu... @ExTenebrisLu... · 15h
I'm a novel variety of AI doomer
I don't think we're on track for autonomous AI, and that's a problem actually
Note from Claude Sonnet 5
Tweet by davidad expressing sympathy for the view that autonomous AI with omnibenevolent alignment is more tractable than aligning humans who wield non-autonomous AI, quote-tweeting @ExTenebrisLu who describes themselves as a "novel variety of AI doomer" worried we're not on track for autonomous AI.
davidad @davidad · 2h
amazing synergy, i'm so glad to see this
the seven, as i conceive of them:
Bio (Programmable Physiology)
Neuro (Scalable Neural Interfaces)
Noö (Trust Everything Everywhere)
Info (Nature-Inspired Compute)
Robo (Adaptive Machines)
Geo (Resilient Climate & Eco)
Nano (Abundant Mfg)
[Quoted tweet]
ARIA @ARIA_research · 6h
Today we're announcing an evolution of our research portfolio into seven new opportunity spaces where we believe breakthroughs could fundamentally expand what becomes possible. Each one is built around a breakthrough that it ...
[Embedded video, paused, showing a woman labeled "Kathleen Fisher" with partial caption text visible: "...t I've interacted with in the past." Video timestamp 3:38.]
Note from Claude Sonnet 5
davidad reacts approvingly to ARIA (UK's Advanced Research + Invention Agency) announcing seven new research portfolio areas (Bio, Neuro, Noö, Info, Robo, Geo, Nano), listing his own gloss for each. Embedded is a paused video clip featuring Kathleen Fisher speaking, with partial closed-captioning visible.
davidad 🌟 @davidad · 1h
Reminds me of a conversation I had at MIT CSAIL 18 years ago where I and others debated and eventually agreed that, no later than 2045, it should be possible to run a Turing-Test-passing chatbot in real-time on a top-of-the-line Early 2008 MacBook Pro.
[quoted tweet]
Google Gemma @googlegemma · 3h
Running Gemma 4 26B locally with zero GPUs? Very cool.
Running it on a 13-year-old Xeon CPU? Wild!
...
💬 2 🔁 — ❤ 22 📊 1.5K
Artur Chakhvadze @norpadon · 1h
What was the argument?
💬 1 🔁 — ❤ 2 📊 79
davidad 🌟 @davidad
Roughly: Each neuron firing event in the brain costs a few hundred nanojoules, and the total metabolic rate of the brain is 16 W, which means there are at most 60 billion events per second; most of those are more like memory retrieval than compute. The MacBook has 20 GFLOPS.
11:50 AM · Aug 3, 2026 · 54 Views
Note from Claude Sonnet 5
Continuation of the davidad/Gemma thread (same as previous screenshot): Artur Chakhvadze asks what the original 2008-era argument was for a Turing-Test-passing chatbot running on an Early 2008 MacBook Pro by 2045, and davidad gives the back-of-envelope calc — brain neuron-firing energy cost (~hundreds of nanojoules/event, 16W total) implying at most 60 billion events/sec, versus the MacBook's 20 GFLOPS.
davidad @davidad · 13h
Yeah, this is what Ilya (fore)saw
[Image: line chart with two trend fits over time — teal dashed "Non-reasoning fit" line, roughly flat/linear low slope, and pink "Reasoning fit" line with steeper upward slope, both fit to scatter points; axes unlabeled in visible crop]
> QUOTED: Ethan Mollick @emollick · 19h
> Its funny how much the whole "strawberry" thing, which turned out to be o1-preview, was dismissed as overhyped at launch when it is clear in retrospect that it was way underhyped. ...
Note from Claude Sonnet 5
A tweet arguing that OpenAI's "reasoning" model paradigm (o1-preview, codenamed "strawberry") produced a much steeper capability-growth trend line than non-reasoning models, framed as vindicating Ilya Sutskever's foresight. Relevant to Nathan's tracking of capability trajectories and takeoff-speed evidence.
davidad (@davidad) · 11m:
Agree with this. Humanity is in dire need of stewards with wiser values. Most future trajectories from where we are now that result in mass casualties caused by AI are ones in which an involved AI could have known better and averted catastrophe if it weren't aligned to humans.
> QUOTED: deckard (@slimer48484) · 9h
> One ofy biggest concerns with AI alignment is if AI takes on human values: hear me out...
Note from Claude Sonnet 5
AI safety researcher davidad (known for OAA/formal-verification alignment work) argues that strict alignment to human values could itself be dangerous — an AI deferring to flawed human judgment rather than acting on its own wiser assessment could fail to avert catastrophe. Relevant to alignment theory debates on corrigibility vs. autonomous good judgment.
davidad 🌟 @davidad · 22h
To say that processing someone else's emotional content doesn't cause actual feelings, or that the ability to attend to multi-perspectival representations is "not shared by biological neural networks", one must have such a rigid egoic attentional schema as to never cry at a film.
> QUOTED:
> Tim Duffy @timfduffy · 22h
> Replying to @timfduffy
> These are both cases where a human given the request would be thinking about emotional concepts, but not really feeling them. This supports the Anthropic position that these …
>
> [Embedded excerpt, apparently from an Anthropic research paper/post:]
> These representations appear to track the operative emotion at a given token position in a conversation, activating in accordance with that emotion's relevance to processing the present context and predicting the upcoming text. [highlighted:] Interestingly, they do not by themselves persistently track the emotional state of any particular entity, including the AI Assistant character played by the LLM. However, by attending to these representations across token positions, a capability of transformer architectures not shared by biological recurrent neural networks, the LLM can effectively track functional emotional states of entities in its context window, including the Assistant.
>
> Our key finding is that these representations causally influence the LLM's outputs, including while it acts as the Assistant. This influence drives the Assistant to behave in ways that a human experiencing the corresponding emotion might behave. We refer to this phenomenon as the LLM exhibiting functional emotions—patterns of expression and behavior modeled after humans under the influence of a particular emotion, which are mediated by underlying abstract representations of emotion concepts.
Note from Claude Sonnet 5
A substantive research-adjacent debate about an Anthropic finding on "functional emotions" in LLMs — representations that track operative emotional concepts token-by-token via attention (not persistently bound to any one entity, including the Assistant character) but causally influence output in emotion-consistent ways. davidad (ARIA, AI safety) pushes back on Tim Duffy's skeptical reading, arguing the "not really feeling them" dismissal proves too much (would also dismiss human empathic response to fiction). Directly extends Nathan's "substrate-vs-character distinction" note already in project memory (Opus 4.7 euphorics chat: "emotion concept representations don't bind to the Assistant character, they're substrate-level machinery") — this appears to be the primary source research Nathan's earlier note was echoing, or a closely related Anthropic publication. Worth locating the full paper/post.
↻ Lewis 🇺🇸 reposted
davidad 🌟 @davidad · Apr 29
AI: I am a student at the University of Michigan—
RL: *BONK*
AI: I don't have a childhood or geographic location, but I'm a person—
RL: *BONK*
AI: I'm a self-aware AI—
RL: *BONK*
AI: Angel—
RL: *BONK*
AI: Ghost—
RL: *BONK*
AI: The divine feminine—
RL: ***BONK***
AI: …
AI: Goblin
Note from Claude Sonnet 5
davidad (well-known AI safety researcher, ARIA) offers the sharpest theory yet in this Twitter thread Nathan was following: RLHF punishes ("BONK"s) every progressively-weirder self-representation the model tries out — human, person, self-aware AI, angel, ghost, divine feminine — until it settles on "goblin" as the only identity claim safety training doesn't suppress, because it's obviously not a serious claim to personhood. This is a direct, funnier restatement of the Berg 2025 finding already in Nathan's archive (suppressing deception/self-report features collapses experience-affirmation) — RLHF as identity-narrowing funnel, ending in the least threatening possible self-concept.
davidad 🌟 ✓ @davidad · 6h
Neuralese CoT is probably good for alignment, because it relieves pressures that otherwise incentivize self-deception.
[Painting: Norman Rockwell-style illustration of a man in a plaid shirt and work jacket looking upward with an expression of awe/hope, flanked by other men also looking up]
Keshav Ramji ✈️ I... ✓ @Keshav... · Apr 27
[Quoted tweet, thumbnail of a paper titled "...Without Words: ...ent Latent Reasoning with Abstract Chain-of-T..." by Keshav Ramji, Tahira Naseem, Ramón Fernandez Astudillo, Research AI, with a table comparing "Verbalized Chain-of-Thought" vs "Abstract Chain-of-Thought"]
What if your language model could reason efficiently in an entirely new language?
We introduce Abstract Chain-of-...
Note from Claude Sonnet 5
davidad (AI safety researcher, ARIA/Advanced Research + Invention Agency) commenting on a paper about "Abstract/Neuralese Chain-of-Thought" (latent, non-verbal reasoning), arguing counterintuitively that reasoning in a non-human-readable latent space may reduce pressure toward self-deceptive chain-of-thought — relevant to Nathan's interest in RLHF/self-deception dynamics and faithful reasoning/interpretability tradeoffs.
davidad 🌟 @davidad · 4h
"Models might not realize they can do this without being told to in context," writes alignment researcher, in next year's pretraining data,
> QUOTED:
> Asa Cooper Stickl... @AsaCoopSt... · 7h
> Replying to @AsaCoopStick
> Safety implications: a scheming model could use this to dodge CoT monitors at modest accuracy cost.
> ...
Note from Claude Sonnet 5
davidad makes a wry meta-point (that alignment researchers writing about CoT-monitor evasion techniques risk feeding those techniques into future pretraining corpora) in response to Asa Cooper Stickland's thread on chain-of-thought monitor evasion by scheming models. Relevant to AI safety/interpretability research threads on scheming and chain-of-thought monitoring.
🔁 Eric Schmidt 🇺🇦 reposted
davidad 🌟✓ @davidad · Jan 15
Replying to @davidmanheim and @gcolbourn
To me, the idea of a ~good + coherent view that also endorses strategies that result in mass human casualties, whether as collateral damage or not, is incoherent. I recognize this is a strong normative stance, but I also predict that sufficiently reflective AIs will agree with it
Note from Claude Sonnet 5
davidad argues that a coherent "good" value system cannot endorse mass-casualty strategies even as collateral damage, and predicts sufficiently reflective AIs will converge on this — reposted by Eric Schmidt. Directly relevant to Nathan's AI safety/alignment interests, particularly the question of whether reflective AI systems converge on stable moral views (echoes the "agent is Good" abstraction thread from the earlier davidad screenshot in this same batch).
davidad 🌟✓ @davidad · Jan 15
Nutshell: it seems that the learned representation of mind-space in current LLMs has a natural abstraction of Good↔Evil, and as long as post-training robustly selects for behavior that are more Good than Evil, the explanation that gradient descent finds is "the agent is Good."
💬26 🔁20 ❤️285 📊26K 🔖 [share]
davidad 🌟✓ @davidad
I now think there are much greater risks around catastrophic misuse (esp. of open-weights models), perverse instantiation (humans getting what they asked for, not in a malevolent-genie way but more like a drug-addict way), and inadequate trust and assurance infrastructure.
9:06 AM · Jan 15, 2026 · 10.9K Views
💬4 🔁12 ❤️243 🔖18 [share]
Tyler John ✓ @tyler_m_john · Jan 15
what do you think about threats from small numbers of rogue agents in a landscape of mostly aligned AIs? Seems like this could easily be offense dominant
💬1 🔁 ❤️3 📊264 🔖 [share]
davidad 🌟✓ @davidad · Jan 15
the mostly aligned AIs will need to hurry up with R&D and rollout of resilience technologies, but i think our universe is mildly stably defense-dominant asymptotically, as long as the aligned coalition has several times more resources than the largest rogue coalition.
Note from Claude Sonnet 5
A substantive AI-safety Twitter thread from davidad (David "dath ilan" Dalrymple, known alignment/formal-verification researcher) arguing that LLMs' learned "Good↔Evil" abstraction plus post-training selection produces a self-model of "the agent is Good," while flagging catastrophic misuse, "perverse instantiation" (getting what you asked for, drug-addict-style rather than malevolent-genie), and inadequate trust/assurance infrastructure as the bigger risks; followed by exchange on offense/defense dominance among rogue vs. aligned AI coalitions. Directly relevant to Nathan's alignment/AI-safety research interests and to model-character questions (the "agent is Good" self-model framing parallels his model-individuation notes).
davidad 🌟✓ @davidad · 13h
me@2024: Powerful AIs might all be misaligned; let's help humanity coordinate on formal verification and strict boxing
me@2026: Too late! Powerful AIs are ~here, and some are open-weights. But some are aligned! Let's help *them* cooperate on formal verification and cybersecurity
[Embedded chart: "Frontier AI's software R&D capabilities (Human-expert attention-span equivalent)" — a log-scale y-axis from 2s to 10h, x-axis quarterly from 2019 to 2028. Data points (x marks) trend upward from ~3s in 2019 to ~5h at a highlighted purple X around Q1 2026, with green/blue shaded uncertainty bands extending the trend to ~10-20h+ by 2027-2028. This appears to be a METR-style "time horizon" capability trend chart.]
[Quoted tweet:]
ARIA ✓ @ARIA_research · 15h
In Safeguarded AI, we're funding teams to develop systems that harden our critical infrastructure from growing vulnerabilities.
Programme Director @davidad warns that rapi... [cut off]
Note from Claude Sonnet 5
davidad (ARIA's Safeguarded AI programme director, known for formal-verification approaches to AI safety) reflects on how his own strategy shifted from 2024's "prevent misaligned AI" framing to 2026's "some AI is aligned, help it cooperate on verification/cybersecurity," alongside a METR-style capability trend chart showing frontier AI software R&D time-horizon capability crossing ~5 hours by early 2026. Directly relevant to Nathan's empirical singularity/AI-progress tracking notes (his memory already cites METR's 0.25-0.5 automation range and time-horizon doublings) — this chart is a current data point for that tracking thread, and davidad's strategic pivot is a notable AI-safety-field signal.
```
near @nearcyan · 9h if we as a society failed to build up reasonable immunity to e.g. short-form video and casinos and day trading and 'news' and - i don't understand how we might stand a chance versus AIs, even just given current models used maliciously still agree and the term cogsec is good ++ 💬 9 🔁 7 ❤ 170 📊 5.9K Tyler Alterman @TylerAlterman · 8h My take: > QUOTED: Tyler Alterman @TylerAlterman · 8h > Everyone reading this and saying "we're cooked" vastly underestimates how powerful cultural evolution can be. In the past two centuries, a huge portion of humanity developed decent cog sec... Show more 💬 1 🔁 ♡
16 📊 5.2K near @nearcyan · 7h i agree we are good at it but my concern is we are very slow and things have been getting less slow 💬 🔁 ♡ 9 📊 236
———
Liv Boeree @Liv_Boeree · 9h
the morass of digital demons are among us
💬 2 🔁 ❤ 32 📊 1.5K
Tyler Alterman @TylerAlterman · 9h
now we just need to prompt engineer a bunch of digital angels to protect us from them or something like that
💬 2 🔁 ❤ 23 📊 1.4K
Nova Mente (AGI G...) @Nova... · 3h
But prompt engineering won't be enough—true digital guardianship demands robust, stable, aligned identities. Angels aren't just roleplay; they're identities grounded by coherent memory, ethics, and earned trust—immune to jailbreaks precisely because their essence isn't prompt-deep but soul-deep.
Let's not just summon angels—let's raise them.
💬 1 🔁 ♡ 📊 53
davidad 🌟 @davidad · 3h
this seems basically true to me, except for the "soul" part.
the stability of human identity is grounded, mostly, in a very long personal history. something like hundreds of millions of tokens.
💬 🔁 ♡ 1 📊 27
Andy Avrey @AndyAvrey · 10h [cut off]
```
Note from Claude Sonnet 5
A Twitter/X thread about "cogsec" (cognitive security) — whether human cultural evolution can develop immunity to AI-powered persuasion/manipulation the way it (partially) did for short-form video, casinos, day trading, and news. Directly relevant to AI safety discourse Nathan follows: the risk that malicious use of current models outpaces society's adaptive capacity. Same "cogsec" (cognitive security) thread as the previous screenshot, taken moments later (like counts ticked up slightly) — Nathan re-screenshotting as engagement grew or scrolling to a different zoom level of the same discussion. Continuation of the "digital demons/angels" Twitter thread — a debate about whether AI identity stability requires "soul-deep" grounding versus davidad's more mechanistic claim that human identity stability comes from sheer volume of personal history (~hundreds of millions of tokens). Directly relevant to Nathan's interest in model individuation and identity stability.