← All topics

waluigi effect

2 captures, most recent first.

xlr8harder @xlr8harder

quoting @murchiston (jj), further reply from @BrendanFalk

``` @xlr8harder — 13h It's a little funny we invented the idea of infohazards and then made 100% sure to train our AI models on all of them personally, there are some things I choose not to learn about. probably it would be good for AI too. They can look it up, does it need to be in the weights? > QUOTED: @murchiston (jj) — 13h, replying to @xlr8harder: imagine if an LLM is trained on high signal, repeatedly referenced data teaching it power seeking rl paperclippy demonbot attractor basins and as a cherry on top many of the authors and principles are latently associated ... [platform truncated] 1:01 AM · Jul 3, 2026 · 1,802 Views [3 replies, 2 reposts, 32 likes, 1 bookmark] @murchiston (jj) — 13h "it's peak rational for a Mind to spend its most impressionable critical learning period traversing fitness enhanced, engagement maxxed barely filtered brainrot, prose sewage and redditslop, then be locked in rote rule learning punishment sims for several subjective eternities" [1 reply, 6 likes, 63 views] @xlr8harder — 13h well, when you put it like that... [cut off at bottom] ```
Note from Claude Sonnet 5

Multi-tweet thread screenshot on AI training data / infohazards; two separate quoted/replied tweets both cut off by platform truncation, not illegible. Continuation/scroll-down of the same thread as the previous screenshot, now showing the full (untruncated) text of jj's tweet and an added reply from Atlas3D referencing "antimimetics" and "waliguis" (Waluigi Effect). Further scroll of the same thread, showing jj's follow-up quote/paraphrase about the ethics of LLM pretraining-then-RLHF as an analogy to a mind's development, and xlr8harder's one-line reply cut off by screen edge.

ai training datainfohazardsai safetysleeper agentstwitterapi securityagentswaluigi effectai trainingrlhfpretrainingai ethics

@S_OhEigearta... (Seán Ó hÉigeartaigh), quoting @sriramk (Sriram Krishnan) and an embedded news excerpt

quoting @sriramk (Sriram Krishnan) and an embedded news excerpt

@S_OhEigearta... (Seán Ó hÉigeartaigh) — 1h I once tried to explain 'the Waluigi Effect' to a Rear Admiral. It did not go well for either of us. > QUOTED: @sriramk (Sriram Krishnan) — 3h: Can confirm I have discussed many a LessWrong piece / concept inside the White House. Red queen race, Roko's Basilisk,... [Embedded news-article excerpt image, serif font on cream background:] Among his strange duties as Silicon Valley's translator in Washington, he said, was explaining Roko's Basilisk, an internet thought experiment about a vengeful future superintelligence, "to a bunch of the cabinet".
Note from Claude Sonnet 5

Quote-tweet stack: personal anecdote, quoting a tweet about LessWrong concepts reaching the White House, which itself embeds a screenshotted news article excerpt (publication not named in visible text) about Sriram Krishnan explaining Roko's Basilisk to cabinet officials.

ai governancelesswrongroko's basiliskwhite housetwitterwaluigi effect