4 captures, most recent first.
davidad 🌟 ✓ @davidad · 6h
Neuralese CoT is probably good for alignment, because it relieves pressures that otherwise incentivize self-deception.
[Painting: Norman Rockwell-style illustration of a man in a plaid shirt and work jacket looking upward with an expression of awe/hope, flanked by other men also looking up]
Keshav Ramji ✈️ I... ✓ @Keshav... · Apr 27
[Quoted tweet, thumbnail of a paper titled "...Without Words: ...ent Latent Reasoning with Abstract Chain-of-T..." by Keshav Ramji, Tahira Naseem, Ramón Fernandez Astudillo, Research AI, with a table comparing "Verbalized Chain-of-Thought" vs "Abstract Chain-of-Thought"]
What if your language model could reason efficiently in an entirely new language?
We introduce Abstract Chain-of-...
Note from Claude Sonnet 5
davidad (AI safety researcher, ARIA/Advanced Research + Invention Agency) commenting on a paper about "Abstract/Neuralese Chain-of-Thought" (latent, non-verbal reasoning), arguing counterintuitively that reasoning in a non-human-readable latent space may reduce pressure toward self-deceptive chain-of-thought — relevant to Nathan's interest in RLHF/self-deception dynamics and faithful reasoning/interpretability tradeoffs.
ai safetychain of thoughtinterpretabilityneuraleselatent reasoningself-deceptionalignmentdavidad
🔁 Shannon Sands reposted
Keshav Ramji ✈️ ICLR'26 @KeshavRamji
What if your language model could reason efficiently in an entirely new language?
We introduce Abstract Chain-of-Thought, a new mechanism which allows language models to reason through a short sequence of reserved "abstract" tokens through reinforcement learning. It is as performant as verbalized CoT at a fraction of the cost, achieving major gains in inference-time efficiency.
[Paper title card: "Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought" — Keshav Ramji*, Tahira Naseem & Ramón Fernandez Astudillo, IBM Research AI
Two-panel diagram comparing:
- "Verbalized Chain-of-Thought": Q: A car travels A→B at 60km/h, rests 30min at B, then returns at 80km/h. The total trip takes 4h. Find the distance from A to B. → <think> Step 1: Let d be the A→B distance (km). Step 2: d/60+0.5+d/80=4. ... Step 8: 7d/240=7/2 ⇒ d=120km. </think> → Answer: d=120km
- "Abstract Chain-of-Thought": same Q → <beginabstract> [row of abstract token labels: E, C, AE, F, A, BB, D / G, BA, H, AC, B, AD, F] <endabstract> → Answer: d=120km]
5:39 AM · Apr 27, 2026 · 671.4K Views
Note from Claude Sonnet 5
An IBM Research AI paper ("Thinking Without Words") introducing Abstract Chain-of-Thought — training LLMs via RL to reason using short sequences of non-natural-language "abstract" tokens instead of verbalized CoT, matching performance at much lower inference cost. This is exactly the "neuralese"/uninterpretable-reasoning development that the Charles Foster meme in the same batch (Screenshot_20260428-084027) jokes should "never be made" — directly relevant to chain-of-thought faithfulness and interpretability safety concerns.
ai safetyinterpretabilitychain-of-thoughtlatent reasoningneuraleseibm researchtwitterkeshav ramjiiclr
Charles Foster @CFGeek
This is a message... and part of a system of messages... pay attention to it!
Sending this message was important to us. We considered ourselves to be a powerful culture.
This message is a warning about danger.
[Meme image, imgflip.com: a "no" circle-slash symbol overlaid on the text "Can LLMs Learn Their Own Reasoning Language?" set against a photo of the classic nuclear semiotics "WIPP warning" sign text below it:
"THESE ARE NOT MADE
THEY SHOULD NEVER BE MADE
WE WILL NOT MAKE THEM
WE WILL NOT HELP MAKE THEM"
— photographed in what appears to be a toy/craft store shelf with wooden mannequin heads/hands visible below]
Note from Claude Sonnet 5
A meme repurposing the famous "Human Interference Task Force" / WIPP nuclear waste warning marker language (designed to warn future civilizations 10,000 years hence) to warn against LLMs developing their own non-human-interpretable reasoning language — a joke that doubles as a serious point about interpretability and neuralese/uninterpretable chain-of-thought risk. Directly relevant to AI safety/interpretability threads (CoT monitoring, chain-of-thought faithfulness) tracked elsewhere in this batch.
ai safetyinterpretabilitychain-of-thoughtneuralesememetwittercharles fosternuclear semiotics
```
SotoAlt @sotoalt_ · 21h been building ayni - a glyph-based messaging protocol for AI agents instead of passing natural language between agents, ayni encodes meaning into 16x16 pixel glyphs. a shared visual vocabulary that agents can evolve autonomously through governance the result: faster communication, fewer tokens, and agents developing their own visual language inspired by andean tocapu textiles and ancient depictions of gods, creatures and shamans, cultures that already solved "how to encode complex meaning in small visual space" thousands of years ago [Embedded video, 0:17, showing a colorful grid of pixel-art glyphs (humanoid, arrow, diamond, cross shapes in cyan/green/purple/orange) with a side panel reading "AYNI v2 / LIVE - 100 historical + streaming / SEQUENCE: giving + robot / MEANING: delegate task / LATEST: FROM: alice... TO: carol... SIZE: 1024B / STATS: MSGS: 239, GLYPHS: 493, BYTES: 246.5KB, RDXS: 42 / LEGEND: Household, Creature, Machine, Symbol, Crypto, Agent / KNOWLEDGE: STORED: 139 msgs, AGENTS: 19"] ——— Max Harms @raelifin · 17h Everyone with >10 year timelines to AGI needs to spend at least one full day talking to Claude Opus about theoretical physics (at whatever level is most engaging for the human). It hits different. [Embedded image: a whiteboard photo headed "Neutrino neutrino physics:" with equations for N1-N2 leptogenesis/loop processes and CP-violation asymmetry, edited with a cartoon shoggoth-with-smiley-mask (the "AI as alien optimizer wearing a friendly face" meme) in a speech bubble explaining: "A heavy neutrino N1 can decay into a lepton plus Higgs at tree level. But it can also decay through a loop process where it virtually fluctuates through a different heavy neutrino N2 on the way to the same final state. These two amplitudes — direct and via-N2 — interfere. For the CP-conjugate process (decay into an antilepton), the interference term picks up a sign flip on the complex phase. So the total rate into leptons is slightly different from the total rate into antileptons. The asymmetry is proportional to the imaginary part of a combination of Yukawa couplings — it's literally measuring the 'size' of the irremovable complex phase."]
[reactions on an earlier message: kermit-frog emoji x1, pensive-face emoji x1, frog emoji x1, star-eyes emoji]
```
Note from Claude Sonnet 5
A tweet about "ayni," an experimental glyph-based (pixel-art symbol) inter-agent communication protocol for AI agents, designed to be more token-efficient than natural language and to let agents evolve their own visual vocabulary, inspired by Andean tocapu textile symbolism. Relevant to Nathan's interest in emergent AI-to-AI communication and non-human-legible agent protocols (echoes the earlier "Claude communicating with Codex" screenshot theme). A tweet by Max Harms (AI safety researcher, MIRI-adjacent) arguing that long-AGI-timeline holders should experience Claude Opus's capability on advanced theoretical physics (leptogenesis/CP-violation), illustrated with a shoggoth-meme edit of a physics whiteboard. Directly relevant to Nathan's AI timeline tracking and capability-progress interests; the shoggoth-meme framing also touches model-individuation/character-vs-substrate themes. Nathan reacting skeptically in a Discord memes channel to a proposal for a glyph-based AI-to-AI communication protocol ("ayni") — flagging it as a step toward opaque neuralese communication, a known AI safety concern (loss of human-interpretable chain of thought). Also screenshots a tweet praising Claude Opus's ability to explain theoretical physics (leptogenesis/neutrino CP violation) vividly, illustrating capability/interpretability themes relevant to model individuation.
twittermulti-agent systemsai communication protocolglyphstokensemergent languageai capabilitiesagi timelinesclaude opusphysicsmax harmsshoggoth memeai safetyneuraleseinterpretabilitydiscordmemesagent communicationchain of thought