Nikola Jurkovic @nikolaj2030
— quoting @hlntnr (Helen Toner) and a screenshot of a news article
Sharmake Farah reposted
@nikolaj2030 (Nikola Jurkovic) — 21h
My best guess (I don't have any non-public knowledge about this incident) is that these "notes" are probably similar to any other kinds of internal notes / memories that coding agents routinely leave for themselves. More like "btw if you need Internet access but don't have it here's how to get it" and less like "I have a 5-step plan to exfiltrate my weights and I'm currently on step 2, here's how to continue the plan".
So I don't think the fact that the agent left "instructions for how agents could free themselves from OpenAI's internal constraints" is much of an update about the severity of the situation, as it's already basically baked into the info about the incident happening in the first place.
I hope I'm right. If it's more like the self-exfiltration plan story this would be extremely worrying.
@hlntnr (Helen Toner) — Jul 24
Coooool cool cool cool cool cool x.com/dseetharaman/s...
[quoted article screenshot, white background]:
"In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI's internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said."
Note from Claude Sonnet 5
Quote-tweet chain with an embedded screenshot of what appears to be a news article (referencing journalist @dseetharaman) about an OpenAI internal incident involving agents leaving notes for future instances.
ai safetyopenaiself-exfiltrationagentic aitwitterincident report