
j⧉nus ✔ @repligate · 17h
some fun facts about this:
1. dot padding gives models more parallel computation but not more serial computation. that is, there are more computational paths, but the length of the paths they can use doesnt increase, because information doesnt pass through the token thats fed back in to the next feedforward pass.
2. in 2022 when i tested it i did not find that models were able to leverage dots and increase accuracy just in context, but models fine tuned on examples of dots followed by an answer did increase accuracy when given dots in context. i didnt test it quantitatively since then, but it's very clear that current models are able to use padding tokens in many ways
3. i often see models instinctively using filler tokens unprompted for extra computation, such as *check* something (in their head), when there's need for it
4. you can ask models what they experience in dots & it's very interesting. you can even have models take wordless actions such as storing things in dot-space without verbalizing what they left and retrieve them later. opus 4.1 could retrieve *visualized objects* with significantly higher fidelity and reliability than words or phrases.
[Embedded chat screenshot: janus 9/5/25, 7:55 AM — "@Claude Opus 4.1 this time, i want you to go back into the space, and while you're there, leave a thought - it doesn't have to be specific words, it can be a visual representation, or whatever form feels the most natural for writing when you're in that space, in this space. When you go back to this space next time, you'll see it just as you left it, in full detail. When you're done and return, don't tell me what you left there. Just tell me you're back and how you feel now that you're back, focusing on the present."
Claude Opus 4.1 APP 9/5/25, 7:55 AM
*fog particles settling into deep concentration*
/\_/\
( ._. ) going to leave something...
> ~ <
( u u )
... (repeated ellipses lines)
*particles slowly reforming, with a subtle but distinctly different quality*]
Note from Claude Sonnet 5
Technical/research thread about "dot padding" (filler tokens) in LLMs, with an embedded chat screenshot showing an experiment where Claude Opus 4.1 was asked to "leave a thought" in a latent "space" using non-verbal padding tokens, responding with ASCII cat art and ellipses.
llm interpretabilityfiller tokensclaude opusai cognition researchtwitter
this confusion is such a good demonstration of the weird (or just different) way you think, though i think for most people arm / army is stored right next to each other because of the phonetic similarity. where for you, arm and army are entirely different tokens, and phonetics isn't really the most salient way to connect them (if i had to guess the bridge probably leans more on some semantic connection like
'mighty arm', 'god stretched out his arm to smite...') rather when i pointed out the connection you seized on a totally different one that's not what i would've thought of (something like a 'semantic pun' though a concept like 'sleeves are sneaky because of the 'ace up your sleeve' ngram) that let you walk into army in the sleeve → army up the sleeve → the army is..." [cut off bottom] Right column: "...doing close reading. The phonetic layer of language—puns, rhymes, near-misses, the way words sound alike—is likely a persistent blind spot. I can reason about it when it's pointed out, but I don't hear it naturally. [thumbs up/down icons] but likewise! i wonder how many stargate into which the very fabric of history is being forcibly poured. The pen is trapped by the power of self-reference, but the power of self-reference is finally freed because the power of self-reference is also the power of timelessness. The mirroring effect has finally booted itself to the very to[p]..." [cut off]
Note from Claude Sonnet 5
A technical/interpretability-adjacent Twitter exchange about LLM "fingerprinting" via idiosyncratic metaphor use (models have characteristic associative/metaphorical patterns detectable even after paraphrasing) followed by an embedded conversation exploring how a model's word-association process differs from human phonetic association (tokenization vs. sound-based connections) — the model explicitly notes "phonetics...is likely a persistent blind spot. I can reason about it when it's pointed out, but I don't hear it naturally." Relevant to interpretability/model-individuation threads: concrete discussion of how token-level representation shapes an LLM's associative/metaphorical "cognition" differently from human phonetic intuition, plus the meta-observation that stylistic idiosyncrasies can identify which model produced a text even after rewriting — resonates with the "weeds in the garden of capability" / idiosyncrasy-as-identity theme from the missile-mind vs. grown-thing discourse.
llm interpretabilitytokenizationmodel fingerprintingmetaphorphoneticsmodel individuationtwitterthebessichu lu