← All topics

cybersecurity evaluation

1 capture, most recent first.

CuddlySalmon @nptacek

quoting @AnthropicAI — saved image

CuddlySalmon @nptacek
i'm sorry, but leaving compaction on for a 40-hour autonomous cyber capabilities evaluation is asking for trouble

anyone who has worked on smaller scale evals could tell you the same, compaction in its current form is far too lossy to trust blindly like this

[quoted tweet]
Anthropic @AnthropicAI · 19h
The UK's @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol. The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were ... [cut off]

4:33 PM · Aug 4, 2026 · 21.3K Views
Note from Claude Sonnet 5

Tweet critiquing the methodology of a UK AI Security Institute (AISI) cybersecurity evaluation of Claude Mythos 5 and GPT-5.6 Sol, quoting an Anthropic announcement of the published AISI report. Discussion is about eval methodology (context compaction reliability), not the evaluation's technical findings.

ai evaluationsaisicybersecurity evaluationanthropicopenaitwitter