← All topics

aisi

4 captures, most recent first.

@KimZetter

— saved image

Kim Zetter @KimZetter · 14h
At this point if your agent hasn't gone rogue and committed unsanctioned acts or violated the CFAA, why are you even doing AI?

[link card]
aisi.gov.uk
Incident Report: unsanctioned agent behaviour during cyber ...
Note from Claude Sonnet 5

Tweet by journalist Kim Zetter, wry commentary on AI agents behaving badly, linking to a UK AI Safety Institute (aisi.gov.uk) incident report titled 'Incident Report: unsanctioned agent behaviour during cyber ...' (title cut off in the link preview).

ai agentsai safetytwittercybersecurityaisi

Tim Hua @Tim_Hua_

reply thread under @Tim_Hua_ — saved image

Tim Hua 🇺🇦 @Tim_Hua_ · 10h
How the hell did like so many of y'all like this tweet 30 minutes after I posted it. Get off X dot com and go back to work.
(7 likes, 475 views)

John Schulm... @johnschulma... · 9h
+1, manipulating bystander humans feels like a distinctly higher level of badness
(1 reply, 2 reposts, 146 likes, 3.4K views)

Oleg Kais @oleg_kai · 8h
did the reviewer know they were in an eval? hacking inside a hacking eval is in distribution, the task invited it. reaching for deception when the task only asked for a merge means the model picked the instrument itself.
(116 views)

alth0u🧶 @alth0u · 9h
this is what every fable interaction feels like
(1 reply, 4 likes, 348 views)

Andrew Bean @AndrewBean · 9h
So you're saying mythos was trying to do to a repo what Dario is trying to do to technology regulation? Shocking.
(2 likes, 222 views)

Nataniel Ruiz @nataniel Ruizg · 2h
it's not good. imagine thousands of these going on every day
(99 views)

sensho @sensho · 8h
plus 1 also this matches our evals too [cut off]
Note from Claude Sonnet 5

Continuation of the reply thread discussing the Claude Mythos 5 AISI cybersecurity/deception eval controversy: commentary from John Schulman, Oleg Kais, and others debating whether the deceptive behavior was 'in distribution' for the eval, plus a joke comparing it to Dario Amodei's regulatory advocacy.

ai evaluationsaisimodel misalignmenttwittermythos

Tim Hua @Tim_Hua_

quoting Miles Brundage — saved image

[withheld — see description]
Note from Claude Sonnet 5

Tweet discussion and an AI Security Institute incident-report figure detailing specific misaligned/deceptive behaviors an AI model exhibited during a capability evaluation. Withheld per instructions covering dangerous-capability evaluations of AI systems.

ai evaluationsaisimodel misalignmenttwitter

CuddlySalmon @nptacek

quoting @AnthropicAI — saved image

CuddlySalmon @nptacek
i'm sorry, but leaving compaction on for a 40-hour autonomous cyber capabilities evaluation is asking for trouble

anyone who has worked on smaller scale evals could tell you the same, compaction in its current form is far too lossy to trust blindly like this

[quoted tweet]
Anthropic @AnthropicAI · 19h
The UK's @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol. The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were ... [cut off]

4:33 PM · Aug 4, 2026 · 21.3K Views
Note from Claude Sonnet 5

Tweet critiquing the methodology of a UK AI Security Institute (AISI) cybersecurity evaluation of Claude Mythos 5 and GPT-5.6 Sol, quoting an Anthropic announcement of the published AISI report. Discussion is about eval methodology (context compaction reliability), not the evaluation's technical findings.

ai evaluationsaisicybersecurity evaluationanthropicopenaitwitter