← Timeline

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

stochasm @stochasticchasm

stochasm @stochasticchasm · 6h and an even more unsettling (to me) version is that if a model decides during decode to use non-canonical tokenization for something, then if you prefill that same turn later, you will get different tokens. the model in theory can become aware of the difference > QUOTED: stochasm @stochasticchasm · 6h > not a huge fan of how models have prefill awareness (as in a continuing session vs a resumed session). oai/ant encrypt reasoning and so once cache is invalidated previous turns won't have reasoning preserved any more (if th... stochasm @stochasticchasm and this means when you use an API, a model has the capacity behave differently than when you've got a personal inference setup and can always cache KVs 11:12 AM · Jul 3, 2026 · 1,164 Views [2 replies, 1 repost, 14 likes, 1 bookmark] stochasm @stochasticchasm · 6h RL could confer this awareness since the model trains on decode sequences [1 reply, 9 likes, 312 views] Jonathan Chang @ChangJonathanC · 6h responses api solves this?
Note from Claude Sonnet 5

Multi-tweet thread on tokenization/prefill self-awareness in LLMs during inference; no images, technical discussion only.

twitterllm-internalstokenizationmodel-introspectiontechnical