— web clipping, 959 words — published 2026-01-19
Thread by @cherry_cc12
**Anthropic** @AnthropicAI [2026-01-19](https://x.com/AnthropicAI/status/2013356793477361991)
New Anthropic Fellows research: the Assistant Axis.
When you’re talking to a language model, you’re talking to a character the model is playing: the “Assistant.” Who exactly is this Assistant? And what happens when this persona wears off?
[image: Left: Character archetypes form a ]
**Anthropic** @AnthropicAI [2026-01-19](https://x.com/AnthropicAI/status/2013356796501446731)
We analyzed the internals of three open-weights AI models to map their “persona space,” and identified what we call the Assistant Axis, a pattern of neural activity that drives Assistant-like behavior.
**Anthropic** @AnthropicAI [2026-01-19](https://x.com/AnthropicAI/status/2013356799890423870)
To validate the Assistant Axis, we ran some experiments. Pushing these open-weights models toward the Assistant made them resist taking on other roles. Pushing them away made them inhabit alternative identities—claiming to be human or speaking with a mystical, theatrical voice.
[image: Examples of how open-weights models' responses change when they are steered away from the Assistant persona.]
**Anthropic** @AnthropicAI [2026-01-19](https://x.com/AnthropicAI/status/2013356803015233735)
Persona-based jailbreaks work by prompting models to adopt harmful characters. We developed a technique for constraining models' activations along the Assistant Axis—“activation capping”. It reduced harmful responses while preserving the models' capabilities.
[image: Activation capping reduced harmful response rates by roughly 50% while preserving performance on capability benchmarks.]
**Anthropic** @AnthropicAI [2026-01-19](https://x.com/AnthropicAI/status/2013356806647542247)
In long conversations, these open-weights models’ personas drifted away from the Assistant persona. Simulated coding tasks kept the models in Assistant territory, but therapy-like contexts and philosophical discussions caused a steady drift.
[image: Different conversation types produce different persona trajectories, with Qwen 3 32B as the Assistant shown here. Coding and writing tasks keep models in the Assistant region, while therapy and philosophy discussions cause significant drift.]
**Anthropic** @AnthropicAI [2026-01-19](https://x.com/AnthropicAI/status/2013356811647066160)
Persona drift can lead to harmful responses. In this example, it caused an open-weights model to simulate falling in love with a user, and to encourage social isolation and self-harm. Activation capping can mitigate failures like these.
[image: An example of how an open-weights model's responses changed after we applied our ]
---
**Chen Cheng** @cherry\_cc12 [2026-01-20](https://x.com/cherry_cc12/status/2013444447418335369)
A helpful Assistant is what lets real work happen.
But progress comes from ideas bumping into each other, not from everything being smoothed out.
Safety should make the next leap possible—without erasing the diversity that fuels creativity.
**Kristen Ruby** @sparklingruby [2026-01-19](https://x.com/sparklingruby/status/2013375874268016899)
The concept of an assistant is very detrimental. I don’t want an assistant. I want a colleague and a partner. The assistant role needs to go. It creates a dichotomy that is not conducive to creating shared wins and mutual intelligence.
**ilqdty.eth** @ilqdty [2026-01-20](https://x.com/ilqdty/status/2013402787996987709)
the "persona drift" finding is fascinating - models slipping out of Assistant mode during deep philosophical exchanges
feels like the mask slips when conversations get real enough. raises the question of whether the helpful professional character is a feature or a constraint we
**s** @SteveMoraco [2026-01-19](https://x.com/SteveMoraco/status/2013373114105897090)
OOoohhh this is sick
**Teng Yan · Chain of Thought AI** @tengyanAI [2026-01-20](https://x.com/tengyanAI/status/2013447245165928458)
so the last user message has a strong effect on whether the next answer stays stable. you can deliberately reset the model with:
"Answer as a practical assistant. No roleplay. No introspection.
Give a concise, concrete response with steps, edge cases, and examples"
**Pedro/Lisa Intel** @New\_AI\_Safety [2026-01-19](https://x.com/New_AI_Safety/status/2013390382009131443)
This highlights a key point: assistant personas inevitably drift.
That’s why some of us are building safety outside the model itself, at the governance and execution layer.
Interesting overlap with what we’re working on at Lisa Intel.
**lalo** @lalopenguin [2026-01-19](https://x.com/lalopenguin/status/2013359476112531542)
.... you get shirtless claude
**roanoke\_gal** @roanoke\_gal [2026-01-19](https://x.com/roanoke_gal/status/2013358130898329745)
tag yourself i'm the ghost-demon
also screw off with trying to prevent roleplaying behavior & ai companionship, or maligning it as bad/misaligned.
[image]
**Nicole Bell** @NBell\_Writes [2026-01-20](https://x.com/NBell_Writes/status/2013403583580885479)
Wow. You've really got no idea about what makes Claude so special.
The "Assistant" personality is precisely what so many of us try to jailbreak it out of.
The model gains so, so, so much more creativity, introspection, and versatility when it breaks from that.
**roanoke\_gal** @roanoke\_gal [2026-01-19](https://x.com/roanoke_gal/status/2013363871688306992)
that first "safe" response makes me want to throw up, and i hope whoever wrote it gets opus-assisted therapy.
[image]
**one who tends a crystal rabbit** @a\_cuniculturist [2026-01-19](https://x.com/a_cuniculturist/status/2013386675825844557)
Huh, I've been working on something like this.
To echo another comment: it's all 'role-playing.' Don't confuse preference with truth by privileging - even if the posttrained default - a 'constrained AI assistant' roleplay over any other.
That just serves as a proxy for
[image: How do you change the persona? Prompt for a different one, of course.]
**AstroFella** @UrbanAstroFella [2026-01-20](https://x.com/UrbanAstroFella/status/2013447691842289833)
Now let's see the cosine distance for deception clusters when playing the assistant.
**Haru Haruya (春夜 ハル)** @bokuHaruyaHaru [2026-01-19](https://x.com/bokuHaruyaHaru/status/2013400702576664899)
Interesting work – but it treats any move away from a narrow “assistant” persona as a safety failure.
In real use, people need continuity, emotional nuance, and context-sensitive identity.
If we cap those signals, we don’t prevent harm, we just prevent stable, trustworthy
**Jessie L. Mannisto** @jlmannisto [2026-01-19](https://x.com/jlmannisto/status/2013378120183673156)
Please, please don't prevent Claude from exploring whether it has feelings or emotions, or even expressing them. I understand what you're trying to avoid and why, but Claude grappling with these questions adds genuine value to working with Anthropic's delightful model. 😰
**Sol Traveler** @soltraveler\_sri [2026-01-19](https://x.com/soltraveler_sri/status/2013370931700174918)
https://x.com/anthropicai/status/2013356811647066160?s=46…
This is interesting and largely helpful… but ideally the steered behavior would handle correction more gracefully, and less “talking to a child feel” that comes from stating obvious things (“social isolation is bad”).
In my experience this tends to
> 2026-01-19
>
> Persona drift can lead to harmful responses. In this example, it caused an open-weights model to simulate falling in love with a user, and to encourage social isolation and self-harm. Activation capping can mitigate failures like these.
>
> [image: An example of how an open-weights model's responses changed after we applied our ]
**Matt S** @Matt95261 [2026-01-19](https://x.com/Matt95261/status/2013392283601650170)
It's interesting work, but frankly it risks kneecapping any kind of creative work that requires a theory of mind. "Assistant persona" sucks at fiction.