← All topics

model monitoring

1 capture, most recent first.

Jonas Geiping @jonasgeiping

— saved image

Jonas Geiping @jonasgeiping
Earlier today we release our report about a vulnerability that allowed us to read out the encrypted thinking traces from many frontier models (thread below!):

A few thoughts:

First, there is an immediate privacy concern with publicly posted reasoning traces (which is also why we took time to release the report after the initial disclosure). We were able to decode the thinking of many json traces posted online, and found private info in there.

Ironically, during this investigation we also had entry into HF during the cybersec incident due to a leaked prod key (but did not exercise the key beyond a whoami  ;)).

Another big update from this for me was actually seeing thinking traces at scale from Anthropic and OpenAI, and all the weirdness and 'casual' misalignment they contain (examples below). I do think model monitoring has an uphill battle ahead of us in the coming year.

---

Finally, there is something to be said for just making thinking traces accessible to all users. I do believe that we would achieve a much broader, much more pluralistic form of oversight through broadly accessible thinking, that would make model deployment safer [cut off]
Note from Claude Sonnet 5

Continuation of the Jonas Geiping thread (seq 741) about the vulnerability decoding encrypted frontier-model reasoning traces, discussing privacy risks of leaked reasoning, the ironic Hugging Face access incident, and an argument for making thinking traces broadly accessible for oversight.

ai securitychain of thoughtmodel monitoringtwitterjonas geiping