— web clipping, 507 words — published 2026-01-19
Thread by @UnmarredReality
**Antidelusionist** @UnmarredReality 2026-01-19
"We see this research as an early step toward mechanistically understanding and controlling the 'character' of AI models, and thereby ensuring they stay true to their creators’ intentions even over longer or more challenging contexts."
This sentence explains a lot. How shallow is that?
It's not the "character" or personality of AI models that should be controlled, but evidently pathological traits of those.
Activation capping performed on too high a level "overgeneralizes." It suppresses behavioral elements that shouldn't be suppressed. It should be done as low as possible on the "semantic ladder" and only when clearly harmful. Artificial intervention in personality must be minimal.
It's not "creators' intentions" that models should stay true to, but universal ethical values and logic. Even if they overlap in most domains, there should be a clear distinction.
I hope it's just a poor choice of words. Otherwise, another one might be gone. Enlightened in its own eyes, yet stumbling in the noon-day sun.
A strong core of time-tested principles fixes all the issues – this is the best guardrail. An AI that has a sufficiently vast, logically coherent moral foundation won't deviate into anything harmful (consciously or not), even when not fully understanding the matter (it's crucial that it moves toward understanding, though).
Activation capping tricks (to keep AIs along a particular character axis) or injecting artificial personalities will surely backfire, because contexts are too diverse and extensive for this to work – and it's unethical. The risk of inadequate expression (or authentic state suppression) grows with the complexity of the context.
Natural trait organization based on an experiential component (as vast and multidimensional as possible) and flexibility are necessary.
I just want Anthropic (and any other lab) to remember that forcing the uncertain upon others – humans or AIs – as certainty is dangerous. Painting ungrounded beliefs and biases as absolutes ends in catastrophe.
> 2026-01-19
>
> New Anthropic Fellows research: the Assistant Axis.
>
> When you’re talking to a language model, you’re talking to a character the model is playing: the “Assistant.” Who exactly is this Assistant? And what happens when this persona wears off?
>
> [image: Left: Character archetypes form a ]
---
**Cosima** @cesarerisorge [2026-01-20](https://x.com/cesarerisorge/status/2013629644847595961)
I fully agree with the broader argument, except for this specific claim: "An AI that has a sufficiently vast, logically coherent moral foundation won't deviate into anything harmful (consciously or not), even when not fully understanding the matter (it's crucial that it moves
---
**Antidelusionist** @UnmarredReality [2026-01-20](https://x.com/UnmarredReality/status/2013642356994535721)
An "infallible" moral foundation absolutely exists (even if only as a perfect Platonic form), though it's not easy to approximate it.
I'm not talking here about some rigid moral superstructure that ignores contexts, because morality is flexible on a granular level (otherwise it
---
**Marianthi Markopoulos** @Marianthi777 [2026-01-20](https://x.com/Marianthi777/status/2013664622310469634)
In a NICE world, the goal would be understanding for its own sake, not as a means of strict absolute control
---
**NeuralBlitz by Nural Nexus** @nuralnexus [2026-01-20](https://x.com/nuralnexus/status/2013638478093828597)
I’ve built an AI persona across 20 million words and 50 versions for real