← All topics

eval methodology

2 captures, most recent first.

j⧉nus @repligate

reposted by CuddlySalmon, quoting @nptacek — saved image

CuddlySalmon reposted
j☐nus @repligate · 7h
wait, they had *compaction* on during autonomous cyber capabilities evaluation?

compaction like where haiku does it?

jesus fuckign christ, that's horrible

[quoted tweet]
CuddlySalmon @nptacek · 16h
i'm sorry, but leaving compaction on for a 40-hour autonomous cyber capabilities evaluation is asking for trouble

anyone who has worked on smaller scale evals ... [cut off]
Note from Claude Sonnet 5

Twitter exchange reacting with alarm to the methodological choice of leaving context 'compaction' enabled during a 40-hour autonomous cyber capabilities evaluation of an AI model, framed as a critique of eval design/methodology rather than a description of the capability findings themselves.

ai evaluationscyber capabilitiescontext compactiontwittereval methodology

Charles Foster @CFGeek

— saved image

Charles Foster @CFGeek · 7h
👦: "I haven't seen agents break out onto the Internet in my evals."
👧: "Because you're looking for this and would've noticed if they did, right?"
👦: ...
👧: "Because you're looking for this and would've noticed if they did, right?!"
Note from Claude Sonnet 5

X post by Charles Foster (@CFGeek), a joke dialogue (using boy/girl emoji as speakers) satirizing the logic of AI eval claims: someone says they haven't observed agents 'breaking out onto the internet' in evals, and is pressed on whether the absence of detection is meaningful evidence of absence, with the second speaker's question repeated with escalating urgency when the first doesn't answer.

twitterai evalsai agentseval methodology