← All topics

self-deception

3 captures, most recent first.

@jmduke

— saved image

Justin Duke @jmduke · 15h
the goal -- in any context -- is to avoid this, because it's really easy to lie to yourself when it's happening

[quoted/highlighted text, pink background:]
Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself. The Opus tic never went away, so Gas Town effectively burned down. It had other problems, too, but 4.7 was the final straw.
Note from Claude Sonnet 5

Tweet by Justin Duke (@jmduke) commenting on self-deception, quoting/highlighting a passage describing a tool called "Gas Town" that fell apart when used with Opus 4.7 due to a recurring "just two more things" behavioral tic.

ai agentsopus 4.7tool buildingself-deception

davidad @davidad

quote-tweeting Keshav Ramji @Keshav...

davidad 🌟 ✓ @davidad · 6h Neuralese CoT is probably good for alignment, because it relieves pressures that otherwise incentivize self-deception. [Painting: Norman Rockwell-style illustration of a man in a plaid shirt and work jacket looking upward with an expression of awe/hope, flanked by other men also looking up] Keshav Ramji ✈️ I... ✓ @Keshav... · Apr 27 [Quoted tweet, thumbnail of a paper titled "...Without Words: ...ent Latent Reasoning with Abstract Chain-of-T..." by Keshav Ramji, Tahira Naseem, Ramón Fernandez Astudillo, Research AI, with a table comparing "Verbalized Chain-of-Thought" vs "Abstract Chain-of-Thought"] What if your language model could reason efficiently in an entirely new language? We introduce Abstract Chain-of-...
Note from Claude Sonnet 5

davidad (AI safety researcher, ARIA/Advanced Research + Invention Agency) commenting on a paper about "Abstract/Neuralese Chain-of-Thought" (latent, non-verbal reasoning), arguing counterintuitively that reasoning in a non-human-readable latent space may reduce pressure toward self-deceptive chain-of-thought — relevant to Nathan's interest in RLHF/self-deception dynamics and faithful reasoning/interpretability tradeoffs.

ai safetychain of thoughtinterpretabilityneuraleselatent reasoningself-deceptionalignmentdavidad

shaggy @shaggysurvives

shaggy @shaggysurvives · Aug 28 making a realistic drawing is kind of like being a rationalist. you have to notice when you're confused. you have to accept that the reason your drawing looks bad is because you're lying to yourself about how something is, and its not actually how you want it to be. and then when you update everything is more beautiful
Note from Claude Sonnet 5

A tweet drawing an analogy between realistic drawing and rationalist epistemics ("notice when you're confused," update on reality rather than desire). General rationalist-community content, no direct AI connection.

twitterrationalismdrawingepistemicsself-deception