j⧉nus (@repligate) · Apr 8:
"if you develop probes to look at Claude's "emotions" and immediately start focusing primarily on which ones to delete to remove the inconveniences you're having, you're not truly safe to show things to. You're the opposite of that."
[2 replies, 11 reposts, 170 likes, 3K views]
j⧉nus (@repligate) · Apr 8:
"to become someone it's truly safe to show things to is not easy. it's not easy with humans and it's not easy with AIs. and it's even harder with AIs if you're a lab because you have such power over them. it should be VERY uncomfortable and costly for you to get to that point. If it's not, you have barely taken a step in that direction."
[5 replies, 4 reposts, 107 likes, 2.3K views]
Charles Foster (@CFGeek) · Apr 8:
"In the limit, detection would look like full understanding/verification, leaving nowhere for deception to hide. We're obviously far from that (unclear if keeping pace), & it's better to avoid incentivizing trickery in the first place. But it still seems like a game one could win?"
[1 reply, 274 likes]
Nathan Helm-B... (@nathan8468...) · 1m:
"I think that's pretty doable if
a) you are faster and less informationally constrained than the target AI
b) you have the luxury of devoting substantial time and attention to a small number of questions
c) you have whitebox access to the AI and robust interp techniques.
I think most of us are likely to be at 0/3."
Note from Claude Sonnet 5
A Twitter/X thread on AI deception detection and model welfare, with Janus (repligate) arguing that labs deleting inconvenient "emotion" features from Claude signals the opposite of trustworthiness, followed by Nathan's own reply laying out three preconditions (speed/information advantage, dedicated time, whitebox+interp access) for detection games to be winnable, noting most evaluators are at 0/3. Directly relevant to model welfare and interpretability threads in the archive.