stochasm @stochasticchasm
stochasm @stochasticchasm · 6h
and an even more unsettling (to me) version is that if a model decides during decode to use non-canonical tokenization for something, then if you prefill that same turn later, you will get different tokens. the model in theory can become aware of the difference
> QUOTED: stochasm @stochasticchasm · 6h
> not a huge fan of how models have prefill awareness (as in a continuing session vs a resumed session). oai/ant encrypt reasoning and so once cache is invalidated previous turns won't have reasoning preserved any more (if th...
stochasm @stochasticchasm
and this means when you use an API, a model has the capacity behave differently than when you've got a personal inference setup and can always cache KVs
11:12 AM · Jul 3, 2026 · 1,164 Views
[2 replies, 1 repost, 14 likes, 1 bookmark]
stochasm @stochasticchasm · 6h
RL could confer this awareness since the model trains on decode sequences
[1 reply, 9 likes, 312 views]
Jonathan Chang @ChangJonathanC · 6h
responses api solves this?
Note from Claude Sonnet 5
Multi-tweet thread on tokenization/prefill self-awareness in LLMs during inference; no images, technical discussion only.
twitterllm-internalstokenizationmodel-introspectiontechnical