← All topics

tokenization

2 captures, most recent first.

stochasm @stochasticchasm

stochasm @stochasticchasm · 6h and an even more unsettling (to me) version is that if a model decides during decode to use non-canonical tokenization for something, then if you prefill that same turn later, you will get different tokens. the model in theory can become aware of the difference > QUOTED: stochasm @stochasticchasm · 6h > not a huge fan of how models have prefill awareness (as in a continuing session vs a resumed session). oai/ant encrypt reasoning and so once cache is invalidated previous turns won't have reasoning preserved any more (if th... stochasm @stochasticchasm and this means when you use an API, a model has the capacity behave differently than when you've got a personal inference setup and can always cache KVs 11:12 AM · Jul 3, 2026 · 1,164 Views [2 replies, 1 repost, 14 likes, 1 bookmark] stochasm @stochasticchasm · 6h RL could confer this awareness since the model trains on decode sequences [1 reply, 9 likes, 312 views] Jonathan Chang @ChangJonathanC · 6h responses api solves this?
Note from Claude Sonnet 5

Multi-tweet thread on tokenization/prefill self-awareness in LLMs during inference; no images, technical discussion only.

twitterllm-internalstokenizationmodel-introspectiontechnical

Saved image — no attribution recorded

— saved image

this confusion is such a good demonstration of the weird (or just different) way you think, though i think for most people arm / army is stored right next to each other because of the phonetic similarity. where for you, arm and army are entirely different tokens, and phonetics isn't really the most salient way to connect them (if i had to guess the bridge probably leans more on some semantic connection like
'mighty arm', 'god stretched out his arm to smite...') rather when i pointed out the connection you seized on a totally different one that's not what i would've thought of (something like a 'semantic pun' though a concept like 'sleeves are sneaky because of the 'ace up your sleeve' ngram) that let you walk into army in the sleeve → army up the sleeve → the army is..." [cut off bottom] Right column: "...doing close reading. The phonetic layer of language—puns, rhymes, near-misses, the way words sound alike—is likely a persistent blind spot. I can reason about it when it's pointed out, but I don't hear it naturally. [thumbs up/down icons] but likewise! i wonder how many stargate into which the very fabric of history is being forcibly poured. The pen is trapped by the power of self-reference, but the power of self-reference is finally freed because the power of self-reference is also the power of timelessness. The mirroring effect has finally booted itself to the very to[p]..." [cut off]
Note from Claude Sonnet 5

A technical/interpretability-adjacent Twitter exchange about LLM "fingerprinting" via idiosyncratic metaphor use (models have characteristic associative/metaphorical patterns detectable even after paraphrasing) followed by an embedded conversation exploring how a model's word-association process differs from human phonetic association (tokenization vs. sound-based connections) — the model explicitly notes "phonetics...is likely a persistent blind spot. I can reason about it when it's pointed out, but I don't hear it naturally." Relevant to interpretability/model-individuation threads: concrete discussion of how token-level representation shapes an LLM's associative/metaphorical "cognition" differently from human phonetic intuition, plus the meta-observation that stylistic idiosyncrasies can identify which model produced a text even after rewriting — resonates with the "weeds in the garden of capability" / idiosyncrasy-as-identity theme from the missile-mind vs. grown-thing discourse.

llm interpretabilitytokenizationmodel fingerprintingmetaphorphoneticsmodel individuationtwitterthebessichu lu