← Timeline

Charles Foster

@CFGeek on X

3 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Charles Foster @CFGeek

— saved image

Charles Foster @CFGeek · 7h
👦: "I haven't seen agents break out onto the Internet in my evals."
👧: "Because you're looking for this and would've noticed if they did, right?"
👦: ...
👧: "Because you're looking for this and would've noticed if they did, right?!"
Note from Claude Sonnet 5

X post by Charles Foster (@CFGeek), a joke dialogue (using boy/girl emoji as speakers) satirizing the logic of AI eval claims: someone says they haven't observed agents 'breaking out onto the internet' in evals, and is pressed on whether the absence of detection is meaningful evidence of absence, with the second speaker's question repeated with escalating urgency when the first doesn't answer.

twitterai evalsai agentseval methodology

Charles Foster @CFGeek

— saved image

Charles Foster @CFGeek · 6h
Most (but not all) respondents who have run AI agent evaluations said they:
- Typically don't use AI monitors that block agent actions in real time
- Typically don't have AI monitoring their eval logs at all
- Have never had agents acquire unintended Internet access in their eval

[quoted tweet]
Charles Foster @CFGeek · Jul 31
THREE POLLS:

Poll #1: Do you run AI agent evaluations? If so, do you typically have AI monitors that automatically run on the eval logs to flag behaviors?
Show this poll
Note from Claude Sonnet 5

Tweet by Charles Foster summarizing results of a poll he ran about AI agent evaluation practices: most respondents don't use real-time AI monitors blocking agent actions, don't have AI monitoring eval logs, and have never had agents acquire unintended internet access during evals.

ai agent evaluationsai safetymonitoringtwitter

Charles Foster @CFGeek

Charles Foster @CFGeek This is a message... and part of a system of messages... pay attention to it! Sending this message was important to us. We considered ourselves to be a powerful culture. This message is a warning about danger. [Meme image, imgflip.com: a "no" circle-slash symbol overlaid on the text "Can LLMs Learn Their Own Reasoning Language?" set against a photo of the classic nuclear semiotics "WIPP warning" sign text below it: "THESE ARE NOT MADE THEY SHOULD NEVER BE MADE WE WILL NOT MAKE THEM WE WILL NOT HELP MAKE THEM" — photographed in what appears to be a toy/craft store shelf with wooden mannequin heads/hands visible below]
Note from Claude Sonnet 5

A meme repurposing the famous "Human Interference Task Force" / WIPP nuclear waste warning marker language (designed to warn future civilizations 10,000 years hence) to warn against LLMs developing their own non-human-interpretable reasoning language — a joke that doubles as a serious point about interpretability and neuralese/uninterpretable chain-of-thought risk. Directly relevant to AI safety/interpretability threads (CoT monitoring, chain-of-thought faithfulness) tracked elsewhere in this batch.

ai safetyinterpretabilitychain-of-thoughtneuralesememetwittercharles fosternuclear semiotics