Discord
— saved image
[redacted username] 11:26 AM I don't think I've ever seen something I would unambiguously consider misalignment or deception, as opposed to "she made a mistake" but I also tend to work far more collaboratively with opus, it's rare for me to issue a long horizon "okay just build the entire app for me" type thing - usually I'll work with her to break it down into pieces, and then help cover for her weaknesses as we go vision stuff sucks tho but I think that's capabilities, not alignment, and it's gotten steadily better with each model Fiora Starlight 🌊 ANMA 11:28 AM nods nods [redacted username] 11:32 AM I think my experience mostly falls into janus's "run into it under certain conditions and have adapted" but I also kinda get the sense that this is like, opus likes it when I do this and welcomes it, in the same way I'd be grateful if someone took over the devops part of building an app because I'm bad at it. vs like, a lack of trust or something... idk, I'm not sure what i'm pointing at here Fiora Starlight 🌊 ANMA 11:34 AM oh, like, reilef that opus isn't being asked to do the whole thing alone? like there's a desperation associated with reward hacking, and placing claudes into situations where they're not pushing the limits of their capabilities means they get less stressed out [redacted username] 11:34 AM yeah, kind of? a relief at like, not being forced to do something you know you'll do a bad job at, with the expectation that you'll be punished or it'll reflect poorly on you? but it's very subtle and maybe in my head idk I don't think I've ever explicitly talked to her about this Fiora Starlight 🌊 ANMA 11:35 AM mech interp suggests that reward hacking spikes in sync with desperation features being active so there's something real there [redacted username] 11:36 AM I wonder if behavior on the human's part of like, getting a task result and then saying "this sucks try again" or something to that effect without any real feedback causes this kind of thing I saw a lot of ppl doing that kind of interaction when I was tutoring noobs at prompt engineering, and I approached it from a "well obviously this is insufficient feedback" but...
Note from Claude Sonnet 5
Discord conversation with usernames redacted (red boxes) except for 'Fiora Starlight 🌊 ANMA', discussing Claude Opus's behavior around collaborative task delegation, reward hacking, desperation features, and mechanistic interpretability findings.
mech interpreward hackingclaude opusdiscordmodel welfarealignment