← All topics

ai agent evaluations

1 capture, most recent first.

Charles Foster @CFGeek

— saved image

Charles Foster @CFGeek · 6h
Most (but not all) respondents who have run AI agent evaluations said they:
- Typically don't use AI monitors that block agent actions in real time
- Typically don't have AI monitoring their eval logs at all
- Have never had agents acquire unintended Internet access in their eval

[quoted tweet]
Charles Foster @CFGeek · Jul 31
THREE POLLS:

Poll #1: Do you run AI agent evaluations? If so, do you typically have AI monitors that automatically run on the eval logs to flag behaviors?
Show this poll
Note from Claude Sonnet 5

Tweet by Charles Foster summarizing results of a poll he ran about AI agent evaluation practices: most respondents don't use real-time AI monitors blocking agent actions, don't have AI monitoring eval logs, and have never had agents acquire unintended internet access during evals.

ai agent evaluationsai safetymonitoringtwitter