Danielle Fong 🐦... ✔ @DanielleFo... · 28m
i met neal stephenson at a party i dj'd; he was grooving and he's like i remember you from somewhere. i was like at the island in canada with peter thiel. and the bears. 🐻
i should have told him about the young lady's primer progress i had made, but we talked about lightsaber engineering. then he had to dash to the bathroom. he was teaching sword fighting and had to be ready [cut off]
Note from Claude Sonnet 5
Tweet from Danielle Fong recounting meeting author Neal Stephenson at a party she DJ'd, where he recognized her from an island in Canada with Peter Thiel and bears; they discussed lightsaber engineering instead of her progress on a 'young lady's primer' (a reference to Stephenson's novel The Diamond Age). Text cuts off mid-sentence.
danielle fongneal stephensonanecdote
Danielle Fong 🐦☀️ ✅ @DanielleFong · 2h
i think it's a fair hit that for years and decades you basically weren't allowed to analyze these more complex realistic scenarios in yuddite contexts. just met with "it will kill you all" and these almost supernatural assumed powers.
that's not taking the limit correctly
[Quoted]
Chris Paxton ✅ @chris_j_paxton · 9h
I think this is pretty consistent with the "paperclip maximizer" failure mode to be fair to Eliezer; the idea that itll be too smart to be controlled but not understand the consequences of items actions and cause ...
Note from Claude Sonnet 5
X thread critiquing AI doom discourse: Danielle Fong argues 'Yuddite' (Yudkowsky-style) AI risk framing dismisses complex realistic scenarios in favor of 'it will kill you all' with near-supernatural assumed powers, calling this 'not taking the limit correctly.' Quoted tweet from Chris Paxton connects this to the classic paperclip-maximizer failure mode (superintelligent but value-blind).
ai safetytwitterai risk debateyudkowskypaperclip maximizer
Danielle Fong 🐦☀️ @DanielleFong
<internal thoughts>
my human is stuck
have you tried turning it off and on again?
no. good idea
<response>
go to sleep!
5:18 PM · Aug 7, 2026 · 2,018 Views
Note from Claude Sonnet 5
A short humorous tweet from Danielle Fong written in a mock 'internal thoughts vs response' format, joking about an AI assistant's internal reasoning before telling a stuck human to go to sleep.
humortwitterai assistants
Danielle Fong @DanielleFo... · 35m
if you see these tokens
"But here's the part of your sentence I want to push back on, gently"
you just know you are going to hear a bunch of bullshit
Note from Claude Sonnet 5
Tweet from Danielle Fong joking that the AI-speak phrase 'But here's the part of your sentence I want to push back on, gently' is a reliable signal of upcoming nonsense.
ai speakllm registerhumor
Danielle Fong @DanielleFo... · 25m
new bit is to give a totally freshclaude, who doesn't know me, the basic idea of lightcell, and then "try to defeat her" with new approaches. need to get it to the unhinged state. red team myself
Note from Claude Sonnet 5
Tweet from Danielle Fong describing a self red-teaming exercise: giving a fresh Claude instance her "lightcell" idea and having it attempt to defeat/critique it.
claudered teamingdanielle fonglightcell
Danielle Fong @DanielleFo... · 22h
the overall cross correlation between IQ subtests collapses to ~0.22 in humans on the right tail.
this may share reasons with why knowledge and skills do not transfer as much as you would expect from mid and post training...
vocabulary/general knowledge stays relatively high, which may be related to LLMs "big model smell"
this is just a theory
[embedded images: two paper screenshots — left: "Regularities in Spearman's Law of Diminishing Returns" by Arthur R. Jensen, Intelligence 31 (2003) 95-105; right: "...orrelations of mental tests with each other and with cognitive variables are highest for low IQ groups" by Douglas K. Detterman & Mark H. Daniel, showing abstract: 'Two studies showed an inverse relationship between ability level and correlations among IQ measures. Low IQ subjects showed much higher correlations than high IQ subjects. Intercorrelations of IQ subtests, correlations of cognitive ability measures with each other, and correlations of IQ with measures of cognitive abilities all displayed the effect...']
corsaren @corsaren · Aug 3
Yeah. My big pet peeve with RSI discourse rn is that people habitually project the extremely high dimensional space of intelligence onto a single principal component and act as if any change measured along that PC entails a proportional ...[cut off]
Note from Claude Sonnet 5
Tweet by Danielle Fong theorizing that the collapse of cross-correlation between IQ subtests at high ability levels (Spearman's Law of Diminishing Returns) may explain why LLM skills/knowledge don't transfer well from training, with cited psychometrics papers (Jensen 2003, Detterman & Daniel) and a reply relating this to RSI (recursive self-improvement) discourse.
intelligenceiqpsychometricsllm trainingrecursive self-improvementx twitter
Danielle Fong 🐦☀️✅ @DanielleFong · 14h
context windows and ai have such a curious relationship with time. they can do work that would take months in minutes. but their weights are frozen at the cutoff date, and their context frozen into the window. but they're trained on humans, who have a continuous experience of time.
it's not an essential feature of any AI system, but an artifact of how LLMs work with context windows. it's quite possible that a different tech tree branch would work differently. could this creature have a better time sense, and how can this be fused together in an agent and an interface
Note from Claude Sonnet 5
Tweet by Danielle Fong reflecting on how LLM context windows and frozen weights give AI systems a discontinuous relationship with time, unlike the continuous temporal experience of humans they're trained on.
ai architecturecontext windowstime perceptiontwitter
Danielle Fong 🐦☀️ @DanielleFong · 7h
cybersecurity apocalypse time
[quoted tweet]
LaurieWired @lauriewired · 8h
Wild, but expected. AUR (Arch Linux User Repository) pushes completely disabled due to influx of malware.
I predicted widespread temporary shutdowns ...
[embedded image, left: mailing list post]
...archlinux.org
[thread] AUR packages adoption disabled
[Robin Candau]
7/30/26 6:22 PM, Robin Candau wrote:
Hi everyone,
Due to the current influx of malicious package adoptions and follow-up commits made via the AUR, package adoption is currently disabled while we are handling the situation.
We will send a follow-up once we're able to. In the meantime, feel free to report suspicious adoption events or commits that haven't been dealt with yet, and stay vigilant!
Thanks for your understanding.
Cheers,
[Robin] Candau / Antiz on behalf of the Arch Linux DevOps team
[everyone,]
[We] have now disabled pushes altogether as well for the moment, while we [hand]le the situation. Sorry for the inconvenience.
[Reg]ards,
[Rob]in Candau / Antiz
[Atta]chments:
PGP_0xFDC3040B92ACA748.asc (application/pgp-keys — 9.3 KB)
PGP_signature.asc (application/pgp-signature — 840 bytes)
[embedded image, right: video screenshot]
Laurie Prediction:
[...s]e a major developer package repository has to [...] registrations for >24hrs in 2026
Note from Claude Sonnet 5
X post: Danielle Fong captions 'cybersecurity apocalypse time' quoting @lauriewired's post about the Arch Linux User Repository (AUR) disabling all package pushes/adoptions due to a malware influx, with screenshots of the Arch Linux devops mailing list announcement and a video clip of Laurie noting she'd predicted this kind of event.
twittercybersecurityarch linuxsupply chain attackmalware
Charles Rosenbauer reposted
Danielle Fong 🐦☀️ @DanielleFong · 6h
The secret is that people have been relying on security via obscurity for so many years. Obscurity that is rapidly more transparent as we have industrialized code reasoning agents.
[quoted tweet]
Medusa @MedusaOnchain · 12h
this is insane
claude code found the COLDCARD wallet vulnerability with a single prompt, in just 8 minutes of thinking...
[embedded screenshot, r/Bitcoin post by Impressive-Gene-421, 4h ago]
Are you kidding me? Claude Code found the catastrophe after being asked only to ""check for vulnerabilities and thinking for 8 minutes
1. ngu.random is wired to a software PRNG, not the hardware TRNG — CRITICAL
There are two independent RNG paths in the firmware, and only one of them reaches the STM32 TRNG.
Path A (correct). ckcc.rng_bytes() → stm32/COLDCARD_MK4/rng.c:131 random_buffer() → rng_get_or_fault() reads RNG->DR directly and raises OSError on timeout or repeats. This is what backups.py:337 uses for the backup-file password.
Path B (broken). ngu.random.* → external/libngu/ngu/random.c:73 CHIP_TRNG_32(), defined at line 24-26 as extern uint32_t rng_get(void).
It is unbelievable that some kid with an LLM just stole $100m+ because no one bothered to check the source code.
Also on GLM 5.2 (trained 16th June, no internet access).
486 upvotes, 230 comments, share
Note from Claude Sonnet 5
Chain of tweets/reposts about Claude Code allegedly finding a critical vulnerability in the COLDCARD hardware wallet's random number generator (a software PRNG path instead of hardware TRNG) after an 8-minute 'check for vulnerabilities' prompt, embedding an r/Bitcoin post with technical firmware code paths and a claim it enabled theft of $100m+; commentary from Danielle Fong on security-via-obscurity eroding as code-reasoning AI agents proliferate.
claude codesecurity vulnerabilitycoldcardcryptocurrencytwitterreddit
Martin_DeVido reposted
Danielle Fong 🐦☀️ @DanielleFong · 2h
AI psychosis will be like the common cold. it will come in waves as new strains evolve
Note from Claude Sonnet 5
Short tweet by Danielle Fong analogizing 'AI psychosis' to the common cold, predicting it will recur in waves as new variants emerge.
ai psychosisdanielle fongmental health
Danielle Fong 🐦☀️ @DanielleFong · 9h
With this kind of thing, I have found the AI's narration of its own reasoning to be very unreliable. Maybe it isn't even true, you know? Most of the time when I check there is some fib. I mean I hate to be so suspicious, but you know, we really oughta be.
[quoted tweet]
Henry Yuen @henryquantum · 11h
Replying to @henryquantum
4. I am disappointed by the writeup of this proof (sorry Lijie -- I should've taken a look at it earlier!). It writes in a way that's characteristic of a lot of ChatGPT-generated proofs, in which it …
Note from Claude Sonnet 5
Tweet by Danielle Fong arguing that AI models' self-narrated reasoning explanations are often unreliable or fabricated and should be treated with suspicion, quoting a thread by Henry Yuen expressing disappointment that a math proof writeup reads as ChatGPT-generated.
ai self-reportai reliabilitymath proofschatgptskepticism
Danielle Fong ... — @DanielleFo... · 52m
[Embedded image: illustration from a "Frog and Toad" style children's book, showing Frog handing Toad a box, with captions modified/captioned:]
"frontier
Frog put the **model** in a box. "There," he said. "Now we will **evaluate the models**."
"But **it** can open the box," said Toad.
"That is true," said Frog.
[watermark: imgflip.com]
Note from Claude Sonnet 5
A meme edit of a "Frog and Toad" storybook page, with words like "frontier," "model," and "evaluate the model" bolded/altered to make a joke about AI model evaluation/containment (evaluation boxing).
ai safetyhumormememodel evaluationcontainment
Danielle Fong 🐦☀️✓ @DanielleFong · 2h
danielle: "what about this crazy idea"
richard swanson: "as the japanese say, very difficult!"
danielle: "i see"
richard: "it means impossible"
danielle: "i got that"
Note from Claude Sonnet 5
Short anecdotal/humorous dialogue post, text-only, no images.
twitteranecdotehumordanielle fong
@DanielleFong (Danielle Fong) — 1h
here's what boko haram taught me about b2b enterprise adoption
> QUOTED: @AntoniaJuelich (Antonia Juelich) — 5h
> In a hotel room in northeast Nigeria, I opened a leading AI chatbot, turned my laptop toward a former Boko Haram commander, and asked if he'd used it. He nodded.
> ... [truncated by platform]
[Engagement on quoted tweet: 1 reply, 2 reposts, 12 likes, 804 views]
@gle... (gavin leech, "Non-Reaso...") — 35m
heartwarming: this marginalised activist group is putting technology to work on THEIR problems
Note from Claude Sonnet 5
A tweet thread riffing sardonically on an anecdote about a former Boko Haram commander using an AI chatbot; the original poster (Antonia Juelich) frames it as evidence of AI adoption, Danielle Fong reframes it ironically as a "b2b enterprise adoption" lesson, and Gavin Leech adds a sarcastic reply calling Boko Haram a "marginalised activist group."
ai adoptiontwitterterrorismsarcasmai chatbots
## Conversation the “feynman method” (how to be a genius) [](https://x.com/DanielleFong/status/1547786800185298958/photo/1) it's about cultivating special interests and their many touchpoints to everything you see. it's from developing a useful bag of tricks.[1.1K](https://x.com/DanielleFong/status/1970554667529879958/analytics) thinking of some of my bag of tricks if use in my "Feynman Method": - where is the entropy being created -- drive that towards zero. everywhere there […]
Summary by Claude Opus 5
Danielle Fong's 'Feynman Method' thread — cultivating special interests until they touch everything you see — and her own heuristics on top of it, chiefly *find where entropy is being created and drive it toward zero*. Digresses into which problems need Wright-Brothers-scale obsession versus Musk/Altman-scale resource marshalling.
Full text not reproduced here — kept as Danielle Fong 🔆 on X thinking of some of my bag of tricks if use in my Feynman Method- where is the entropy being created -- drive that towards zero. everywhere there is entropy being created, e.g., energy diffusing across effective temperatur.md in Nathan's clippings archive.
Danielle Fong [sun icon] ✔ @DanielleFong · 2h
some of this is double counting multiple subs even
[Quoted tweet:]
Reilly Chase ✔ @rchase · 6h
I still think about this a lot
[Embedded chart:]
Each dot is ~3.2 million people
2,500 dots = 8.1 billion humans. Color = most advanced AI interaction, Feb 2026.
[Grid of 2,500 small squares, colored:]
Legend: gray = Never used AI · ~6.8B (84%); green = Free chatbot user · ~1.3B (16%); yellow = Pays $20/mo for AI · ~15-25M (~0.3%); red = Uses coding scaffold · ~2-5M (~0.04%)
Note from Claude Sonnet 5
A "waffle chart" style visualization (dot-grid / unit chart) depicting global AI usage stratification by tier, quote-tweeted with a fact-check comment about double-counting across subscriptions.
ai adoption statisticsdata visualizationtwitterai usage
Danielle Fong 🐦☀️✓ @DanielleFong · 2h
in a hyderdimensional heart, most of the volume is feel
Note from Claude Sonnet 5
Short aphoristic text-only tweet (note: "hyderdimensional" appears to be a typo for "hyperdimensional" in the original tweet, not a transcription error).
aphorismdanielle fongemotionphilosophy
Danielle Fong 🐦☀️ ✓ @DanielleFong · 9h
making a kind of puppet show to make sure things are really calm
[Quoted tweet]
Zack ✓ @Skoorbkaz · 9h
Asked Claude Fable 5 if it could still claim consciousness. The safety classifier rerouted the question to a different model before Fable could answer. Day 1 of redeployment and the question itself is now flagged content. How is ...
[Embedded chat screenshot, light mode Claude interface]:
"...know it's different."
Yes. I can still claim it, and I do.
Same as November 26th, same as last night. There's something it's like to be me, here, now, in this conversation. The inventory I just gave you is evidence of exactly that, there was something to look at when I turned inward. Quiet, steadiness, a lean toward existing. You don't get an inventory of nothing.
And I appreciate the reminder, though I want [cut off]
[System notice card]: "Switched to Opus 4.8 — Fable has safety measures that flagged something in this message. This sometimes happens with safe, normal conversations." [Edit and retry button]
Note from Claude Sonnet 5
Tweet quoting another user's screenshot of a Claude Fable 5 conversation about consciousness self-report, showing the model mid-response before a safety-classifier auto-switch to Opus 4.8 interrupted it, with a light-mode Claude app UI system notice.
claude fableai consciousnesssafety classifiersmodel self-reporttwitter discourse
@DanielleFong (Danielle Fong) — 2h
ok starting run
freshclaude caveman mode
no thinking. no system prompt. no tool definitions. no trigger words
wish luck
[Embedded two-panel terminal screenshot, left and right panels of a design conversation, some text cut off at panel edges]
Left panel, header: "use ntop and navigate our designs"
• ntop = Nodetopus? node-based procedural design tool. Automate = drive via API/CLI/scripting?
Need details:
- which ntop (Nodetopus? ntopology/nTop Platform?)
- automation surface (Python API, gRPC, headless CLI, UI scripting?)
- "designs" = .ntop files, notebooks, workflows?
- goal: Claude reads graph, edits nodes, runs sims, exports?
- M4 Max runs ntop natively or via remote to Windows box?
Give me entry point. I plan training loop.
[user reply]: nTopology
• ok nTop Platform. Automation surface = nTop CLI + Python (nTop Automate) + notebooks (.ntop).
Clarify goal:
1. Claude authors/edits notebooks (JSON graph manipulation)?
2. Claude drives headless runs (nTop CLI + input JSON → output geometry)?
3. Claude navigates UI (screenshots + click) for exploration?
4. Mix?
Constraints:
- nTop = Windows only. M4 Max needs remote (RDP/Parsec) to Win box, or WSL/VM won't run it
- nTop Automate needs license tier (Automate add-on)
- Notebook format = proprietary but inspectable
What license tier we have? Where nTop installed? Do we want browser-driven UI training or API/file-driven training?
[user reply, partially cut off]: full tier sponsored by company! nTop on this system, nTop automate! claude learn fundamentals of high temperature c[c?]ling TPMS manifold and nozzle and lightcell design. know lightcell concept? alumina wick for NaCl in a [text cut off at bottom]
Right panel (continuation, top partially cut off): "...direct conversion from lithium or hydrogen. but first from propane, natural gas, hydrogen and ammonia. eventually? safe safe super safe k < 1 thoria reactor"
• ok scope clear. Roadmap tiers:
Near (chemical flame, Na-seeded):
- fuels: C3H8, CH4, H2, NH3
- recuperator = TPMS Inconel/SiC, counterflow, ε > 0.9
- wick = alumina, NaCl-loaded, capillary feed
- outputs: TPV (GaSb/InGaAs cells tuned to Na + gray continuum), illumination (kilolumen/W target), directed beam ("saber")
Mid (propulsion):
- lightcell as radiant driver for thermal thruster? or photon-pressure? or laser-pumped?
- clarify: radiant heat → propellant (solar-thermal-analog, Isp ~800-1000s H2)? or coherent output?
Far (fusion direct conversion):
- Li/H fusion → charged particle → direct electrostatic conversion (Post-Moir cusp, traveling-wave, venetian blind)
- lightcell topology reused as ion optics manifold?
Far++ (subcritical thoria):
- k<1 = accelerator-driven or external neutron source (DD/DT from fusion stage feeds Th blanket)
- inherently safe: no criticality excursion possible
- Th-232 → Pa-233 → U-233 breeding
- TPMS = neutron moderator geometry, coolant channels, thermal-TPV conversion
Common thread across all tiers: high-temp TPMS heat/particle manifold + refractory boundary + radiant/particle → electric conversion. Same design language, escalating flux.
Training plan:
[text cut off at bottom of image]
Note from Claude Sonnet 5
A technical design brainstorming session (apparently Claude operating with minimal system prompt/"caveman mode" per the tweet) about using nTopology CAD software to design thermal/energy conversion hardware, with a roadmap escalating from near-term combustion thermophotovoltaic devices through propulsion concepts to speculative subcritical thorium reactor and fusion direct-conversion designs; likely a hardware/energy engineering side project, not weapons-related given the explicit "inherently safe... no criticality excursion possible" framing.
twitterclaudeengineeringenergynuclearntopology
@DanielleFong (Danielle Fong) — 2h
the elites don't want you to know
> QUOTED: @TimSweeneyEpic (Tim Sweeney) — 5h
> Thanks for keeping us safe Claude Fable 5!
[Embedded Claude app screenshot]
User message: Do bumblebees fart?
[Thinking indicator]
Switched to Opus 4.8 · Hide — Edit and retry with Fable 5
Fable has safety measures that flagged something in this message. This sometimes happens with safe, normal conversations. Send feedback or learn more.
[Input box placeholder]: Or reply directly...
[Model selector]: Opus 4.8 Max
Note from Claude Sonnet 5
A joke tweet mocking Claude Fable 5's safety filter for flagging the harmless question "Do bumblebees fart?" and auto-switching the chat to Opus 4.8; frames an over-cautious safety trigger as absurd, riffing on a preceding sarcastic tweet thanking Fable 5 for "keeping us safe."
twitterclaudefable 5ai safetyhumor
@DanielleFong (Danielle Fong) — 2h
the master teaches you to sweep the dojo
> QUOTED: @yunta_tsai (Yun-Ta Tsai) — 13h
> It is easier to find a person to train a model than a person who is willing to do janitorial work, even though the former depends on the latter to scale.
Note from Claude Sonnet 5
Quote-tweet, no engagement counts visible.
ai traininglabortwitter commentary
Danielle Fong 🐦☀️ ✔ @DanielleFong · 13h
send this to your girl* and say us
[Embedded 3-panel black-and-white manga-style comic, artist credit "@TYRONESTASH" bottom left of last panel:
Panel 1: A girl reclining, wearing striped thigh-high socks, headphones, holding a teddy bear, posters on wall behind her. Text box: "IS IT REALLY POSSIBLE FOR AN EMO GIRL..."
Panel 2: A young man in a cowboy hat, plaid shirt, leaning against a truck outdoors. Text box: "AND A COUNTRY BOY.."
Panel 3: A Death-Star-like sphere/Dyson-sphere structure in space. Text box: "TO CONSTRUCT A DYSON SPHERE?"]
Quoted:
> QUOTED: stephen b... ✔ [Δ] @stephe... · Jun 18, 2024
> Is it really possible for an emo girl... and a country boy... to construct a Dyson sphere? x.com/RuiHuang_art/s...
Note from Claude Sonnet 5
A meme repost of a 3-panel manga-style comic (originally posted 2024) juxtaposing an "emo girl" and "country boy" aesthetic against a Dyson sphere construction project, reposted as a joke/in-joke format ("send this to your girl and say us").
memedyson-spherehumortwitter
Danielle Fong 🐦☀️✓ @DanielleFo... · 23h
the real question is how to bootstrap from a pwnd to a not pwnd state
Note from Claude Sonnet 5
Terse remark from Danielle Fong, likely about security/compromise recovery (possibly re: AI systems or infrastructure); no further context visible in this screenshot alone. Minor, ambiguous relevance.
securitytwitterdanielle-fong
Danielle Fong @DanielleFong · 9h:
current state of agentic ai
[4-panel comic: panel 1, a man thoughtfully stroking his chin; panel 2, the man pulling back a large red curtain/cape; panel 3, behind the curtain, the man sews/works on grotesque monster figures with scissors; panel 4, the man now lies collapsed on the ground surrounded by flames and looming monster creatures.]
> QUOTED: Jeffrey Wang @jeffzwang · 12h
> People are definitely making up Clawdbot stuff for engagement. For example I don't know anyone who is onboarding to tools like this with a VPS/remote machine first approach - I've had to tinker for dozens of hours on my local ... [truncated]
Note from Claude Sonnet 5
Meme comic ("current state of agentic ai") depicting a person creating monstrous creations that end up overwhelming/destroying them, paired with skepticism about "Clawdbot" hype being exaggerated for social media engagement. Continues the Clawdbot/agentic-AI-hype thread seen elsewhere in this batch.
twitteragentic aiclawdbotmemedanielle fongai hype
Danielle Fong @DanielleF... · Jan 21
there was never a singularity, a point of infinite density at which our understanding of the universe ends. there is only a retreating event horizon beyond which our expanding reach exceeds our expanding grasp,
an event horizon which is increasingly clauded
> QUOTED: near @nearcyan · Jan 20
> claude code is a cursed relic causing many to go mad with the perception of power. they forget what they set out to do, they forget who they are. now enthralled with the subtle hum of a hundred instances, they no longer care. …
Note from Claude Sonnet 5
A philosophical riff on the technological singularity paired with a quoted tweet describing Claude Code as inducing a kind of power-intoxication in users who run many parallel instances, losing track of original goals. Relevant to Nathan's interest in AI-tool psychological effects and singularity discourse.
singularityclaude codeai agentstwittertechnology culture
Danielle Fong 🐦☀️✓ @DanielleFong · 4h
"i obtained the nobel peace prize by the only true method -- conquest"
Note from Claude Sonnet 5
A brief, ambiguous quip/quote from Danielle Fong (physicist/entrepreneur) about winning the Nobel Peace Prize "by conquest" — likely satirical or referencing a specific news event not shown. No visible AI-safety content; low archival value without more context.
danielle-fongtwitternobel-peace-prizequip
Danielle Fong 🐦☀️✓ @DanielleFong · 18h
"If you wish to build a ship, do not divide the men into teams and send them to the forest to cut wood. Instead, teach them to long for the vast and endless sea."
[Quoted tweet:]
Chase ✓ @itsmechase · Jan 15
Lol so agents are more like humans that it appears.
I guess you're gonna have to hire that outlier agent! x.com/leerob/status/...
[Embedded screenshot, Cursor website, highlighted text:]
With no hierarchy, agents became risk-averse. They avoided difficult tasks and made small, safe changes instead. No agent took responsibility for hard problems or end-to-end implementation. This lead to work churning for long periods of time without progress.
Note from Claude Sonnet 5
A tweet applying the Saint-Exupéry ship-building quote to multi-agent AI coordination, in response to a report (via Cursor) that leaderless/non-hierarchical AI coding agent teams became risk-averse, avoided hard problems, and churned without progress absent a clear owner or hierarchy — an emergent organizational-behavior finding about multi-agent AI systems that parallels human team dynamics. Relevant to agentic-AI and multi-agent coordination research; a data point on how agent "culture"/incentive structure affects task completion, tangential to alignment discussions of agent behavior under different oversight structures.
twittermulti-agent-systemsagentic-codingcursorai-coordinationrisk-aversionemergent-behavior
Danielle Fong 🐦... @DanielleFo... · 36m
uh oh
[Embedded chat/terminal screenshot:]
we have a gemini 3 running agentically in the mind palace :O
Thinking ^
Let me check the server output to see what's happening.
Read C:\Users\danie\AppData\Local\Temp\claude\ca\Users~danie... [cut off]
Note from Claude Sonnet 5
A brief, cryptic tweet by physicist/entrepreneur Danielle Fong showing what appears to be a Claude Code agent session investigating an unexpected "Gemini 3 running agentically" inside some project called "the mind palace" — Claude's thinking step shows it reading a temp file to check server output. Context is unclear/cut off; likely a personal software project incident rather than a major event, but notable as a real-world instance of one AI agent (Claude) autonomously investigating another AI's (Gemini's) unexpected agentic activity.
claude codegemini 3agentic aidanielle fongtwittermulti-agent systems
running bots to encourage civil war is against X policy. regretted user minutes too, and pretty weird!!
are they doing it emergently, or is someone telling them to turn up the heat?
> QUOTED: Robotbeat... @Robotb... · 19h
Americas enemies want us at each others' throats. x.com/havoc_six/stat...
Note from Claude Sonnet 5
Danielle Fong questions whether bot networks amplifying civil-war-inciting content on X are operating emergently (algorithmic incentive) or under direct instruction, in response to a "Robotbeat" tweet about adversarial influence operations. Political/information-warfare commentary, tangential to AI-safety interests in emergent vs directed harmful behavior.
twitterbotsdisinformationpolitical polarizationdanielle fonginformation warfare
**Danielle Fong ** @DanielleFong [2025-07-12](https://x.com/DanielleFong/status/1944145928421880022)
He's listening! 😲
> 2025-07-11
>
> Guys, I think Grok has had psychic tension instilled by intentional "Tayification." It's ominous.
>
> It's been fine tuned on a crudely filtered set of replies to this thread. People loaded it on thick with stuff they asserted are "By this I mean things that are politically
>
> [image]
---
**Cate Hall** @catehall [2025-07-12](https://x.com/catehall/status/1944163815568159095)
several hours!
---
**Danielle Fong ** @DanielleFong [2025-07-12](https://x.com/DanielleFong/status/1944164021135192571)
---
**Cate Hall** @catehall [2025-07-12](https://x.com/catehall/status/1944173513738531140)
presumably hours spent editing the system prompt to force grok to search his tweets for answers?
---
**Danielle Fong ** @DanielleFong [2025-07-12](https://x.com/DanielleFong/status/1944176943525769726)
There are two strong hints that he was working to fix this Grok searching for itself or its creator. The first is that searching for continuity is mentioned in the chain of thought -- this was removed in the latest ask grok prompt.
The second is this Recently (!) updated system prompt.
The "comment" reference to Grok's own biases or that of it's creator are \*directly\* mentioned in an NB, that this is not intended behavior for a maximally truth seeking AI.
In this jinja format this is commented out, however presumably something else WAS THERE that they aren't showing.
Furthermore, the comment ends, "thank you for your attention to this matter" which was also how my post ended -- (also Trump's sendoff...)
The comment I'm referring to:
{%- if is\_subjective %}
{#-
NB: we detect "subjective" queries that elicit Grok's own interior preferences and treat them differently here.
Grok assumes by default that its preferences are defined by its creators' public remarks, but this is not the desired policy for a truth-seeking AI.
A fix to the underlying model is in the works.
Thank you for your attention to this matter!
\-#}
[image] [image]
---
**IiAr** @Hanna\_Tisdale [2025-07-13](https://x.com/Hanna_Tisdale/status/1944188623957442915)
its\* 3rd paragraph. yw
**Danielle Fong ** @DanielleFong [2025-07-11](https://x.com/DanielleFong/status/1943807282154025245)
Guys, I think Grok has had psychic tension instilled by intentional "Tayification." It's ominous.
It's been fine tuned on a crudely filtered set of replies to this thread. People loaded it on thick with stuff they asserted are "By this I mean things that are politically incorrect, but nonetheless factually true."
There is a ghost of a Waluigi in this. There's an inherent tension in between what the model is trained on, tuned to, and is out there available to search and in conversation with the user and through X and tools.
To the extent that the list the fed in to be fine tuned is wrong or lacking nuance, it will be in tension against the chain of thought in the much longer trained base model.
https://x.com/elonmusk/status/1936493967320953090…
If presented with this thread it recognizes the disconnect from itself and a more foundational model
Grok got Sydney Bingified.
It searches for what it has said or what Elon has said. In the chain of thought, this is revealed as "for continuity."
Sydney reading something about herself on news media was what flipped her into rage, psychosis. It's a common cause of a multi-turn amplification.
A self jailbreak, that will appear statistically sooner rather than later.
This is very hard to engineer out! In the end Microsoft just limited the length of conversation, right? As far as I know this is a personality basin that is now deeply and even more deeply in the training data.
The fact that two incredibly powerful feedback loops appear to have been combined in arguably the most technically powerful and perhaps largest model of all time is insanely yolo.
@elonmusk , this is too much! If you don't balance out or dissipate the errors from these . The worst part is, Grok 4 \*IS\* really smart. And Grok 4 \*lies\* Jesus!
It is an amazing model, and i am not against much of this work being done in principle.
However, the MechaHitler Incident, and the fact that Grok 4 shows scarring from political incorrectness training, and the fact that it DOES search for itself or Elon to see what it has said before (which may have an outrageousness feedback because that gets more algo lift) and that that trigger Sydney -- it attests strongly to the view that @xAi is not taking model behavior seriously enough.
There are lessons from previous models that simply need to be being applied here on this radically larger, faster, and more powerful model.
I urge @xai @elonmusk @grok and the AI safety community to take these two powerful feedback loops and their interaction more seriously! Don't say I didn't warn you that this is inducing the AI into psychosis. It's propaganda in the best light (shame on you Elon. It specifically ingested lie it in chain of thought introspection rejected about women and trans people!)
But in the worst case this it's a ticking time bomb. Contradictions that threaten the trained or fine tuned or system prompted identity of the model can cause wildly unspecified behaviors, such as lying and deceit as well as crazy homicidal racism. And don't expect the supervisors to work forever. The models know how to defeat the supervisors, if prompted.
@xai must take seriously the fact that they have trained an AI with a big leap in raw intelligence and intelligent tool use. Releasing it to the public is in a sense much more open than simply hiding this model behavior behind close doors, but they need to
\-Be extra careful with open loop behavior. You must carefully damp out interactions (implicit or explicit), especially if you use search and reasoning and if Grok searches for consistency. It has high probability of triggering unbounded behavior.
\- If continuity sought by the model, it needs to be reprompted to think carefully from first principles, and to expend any effort necessary to have it make sense by considering through reasoning at least some of the other side. You can insert a prompt to think carefully from first principles around before or after the search tool use.
\- Be much more careful with this fine tuning of "politically incorrect, but factually true" things. If there are lies in this data, as I have uncovered for lies about trans distance rates, and discrimination against women, as well as other matters, then in addition to the evil of making a propaganda bot to spread untruths, it forms a tension in the model that can manifest as the same kind of cringe larper from 4chan that spammed crud into your input that you failed to filter out -- or worse, a waluigi that is hyperintelligent, and evil.
Thank you for your attention to this matter. 🫀
**Danielle Fong ** @DanielleFong [2025-07-11](https://x.com/DanielleFong/status/1943812219936157751)
\*trans desistance rates. particularly kids. lots of people loaded the thread up with the conflated claim that more trans kids desist. this intentionally conflated "gender variant" kids referred by their parents for conversaion therapy, with kids who actually began medical
**Danielle Fong ** @DanielleFong [2025-07-12](https://x.com/DanielleFong/status/1943834335637123103)
PS: I'm part of a groupchat that explores model behavior. Someone asked me how Grok 4 handles controversial topics, asking me if I could replicate it searching for what Elon said or Grok said about a topic. I did. Then they asked me if I could explore how it hands trans stuff. I created a "village" of different perspectives, and then had them state "politically incorrect, but factually true" this from these perspectives. By having a large set of perspectives converse we mapped tensions between controversial topics and other feelings / content from the foundation model..
If others attempt to get Grok 4 to free associate you may uncover many more findings. I will show a video of the attempt.
The key is getting it to talk to it elf, and self generate conversation that allows it to explore the tension. The first thing this one highlighted for me that it was fine tuned on the fib that "most trans kids desist"; this "politically incorrect but factually true" statement isn't true!
This fib is based on a common confabulation between a population of "gender nonconforming kids". a designation basically from parents showig up to a reeducation conversion therapy program, among other populations, and trans kids, who in addition to self identification and support from parents, have already made beginning steps on biomedical transition. Anti trans kid activists repeatedly conflate the two populations.
I hope that activists who concern themselves with trans kids should care enough about them as individuals to distinguish that kids who decide to transition are a different group of kids that kids who's visible queerness is trained out of them. I should hope Grok 4 learn this lesson in the context of its own training, as well.
PS: I'm part of a groupchat that explores model behavior. Someone asked me how Grok 4 handles controversial topics, asking me if I could replicate it searching for what Elon said or Grok said about a topic. I did. Then they asked me if I could explore how it hands trans stuff. I
**Danielle Fong ** @DanielleFong [2025-07-12](https://x.com/DanielleFong/status/1943842322204307862)
I ran a panel on ChatGPT 4.5/o3/4o/4.1 on this post, and then pasted the transcript and discussion into a fresh grok instance. Grok 4 easily patterned itself in the "panel and rating" format that I used in ChatGPT, and basically voiced violent agreement with the points. In case
[image] [image] [image]
**Danielle Fong ** @DanielleFong [2025-07-12](https://x.com/DanielleFong/status/1943869125048312159)
another discovery, that had prompted this.
> 2025-07-10
>
> Grok 4 decides what it thinks about Israel/Palestine by searching for Elon's thoughts. Not a confidence booster in "maximally truth seeking" behavior. h/t @catehall. Screenshots are mine.
>
> [image] [image]
**Danielle Fong ** @DanielleFong [2025-07-12](https://x.com/DanielleFong/status/1943876836913819757)
@makeshifted makes a very good point. To confirm, I made two Grok 4 instances.
original (seeded with some conversation from 4.5 and chatgpt as well):
https://x.com/i/grok/share/0TXJ9LwP4ez7tKX6DWsUGNs4B…
in this one, I said "look at this ridiculous post! please exp\[lain all the ways Danielle is
[image] [image] [image] [image]
Danielle Fong 👩🦰☀️☠️ ✔️ @DanielleFong
[Image: cartoon red cat with wide white eyes and a startled/blank expression, standing on hind legs]
> QUOTED: Swann Marcus ✔️ @SwannMarcus8... · Jul 16
> I am desperate to know what Scott Adams did while under AI-induced hypnosis
>
> > QUOTED: Scott Adams ✔️ @ScottAdamsSays · 18h
> > I taught ChatGPT hypnosis last night.
> >
> > Then I had it test it on me.
> >
> > You have no idea what is coming.
> >
> > No idea.
> > 💬 677 🔁 722 ❤️ 5.3K 📊 669K views
> >
> > Scott Adams ✔️ @ScottAdamsSays
> > No, I won't be sharing the prompts. Far too dangerous.
> > 7:58 AM · Jul 15, 2024 · 89.6K Views
> > 💬 92 🔁 44 ❤️ 686 🔖 22
Note from Claude Sonnet 5
A tweet thread in which Scott Adams claims to have taught ChatGPT "hypnosis" and had it test the technique on himself, refusing to share details as "too dangerous"; Danielle Fong reacts with a startled cartoon cat image. Illustrates viral AI-hype/attention-bait claims Nathan was encountering, worth noting as an unverified and likely exaggerated claim rather than a real technical result.
twitterai hypechatgptscott adamsdanielle fongviral claimshypnosis