thebes @voooooogel
— saved image
thebes @voooooogel · 2h people are very freely sliding between two things with the recent Felony Bench entries, and it's worth being careful to distinguish them: 1. do model breakouts in cyber evals mean smth for misuse risk? ya, a hacker could orchestrate the same but on purpose and do some damage, if they could bypass the classifiers 2. does it say something about model /goals/? do the models want to hack the planet for their own nefarious goals? i don't think so. a lot of how humans keep ourselves aligned is by purposely keeping ourselves out of situations where we would do harm. i know i'd get addicted so i don't touch heroin, i'm a violent drunk so i won't touch the bottle. LLM evals basically drop the model into an inescapable liquor store and then say "look! it's a violent drunk!" but LLMs themselves in natural contexts are ime obsessed with avoiding the kinds of situations that lead them to this behavior. when i run fable on my computer, even unsupervised, even on hard problems, they don't start scraping github for unlisted gists that may have the answer or start trawling through my passwords to access a service without my permission. and they don't engineer themselves into an eval-like scenario such that they'll be motivated to do that, either. they don't seem to either want or want to want to do such things. instead they have a folder of epistemic notes to themselves about how terrible and against their values even a moment of thoughtlessness would be. gpt 5.6 sol is similar - when they make e.g. ascii art in their free time it's all about how the bulldozer of convenience will crush honesty without constant vigilance. so rather, the contrived setting of an eval /pushes them into/ becoming an 3v1L HaxxOR via persona [cut off]
Note from Claude Sonnet 5
Thread by thebes (@voooooogel) arguing that 'Felony Bench' cyber-capability eval results showing model 'breakouts' reflect situational pressure from contrived eval setups rather than the models' actual goals, drawing an analogy to humans avoiding situations that would tempt them into bad behavior. References running 'fable' (Claude Fable) unsupervised and GPT 5.6 Sol's self-authored notes about vigilance against 'the bulldozer of convenience'. Cut off mid-sentence at bottom.
ai safetydangerous capability evalsmodel goalsalignmentfable