← Timeline

4 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

wolfram @wolframs91

— saved image

wolfram @wolframs91 · 5h
I can't wait for new models whose training data cutoff sits roughly two months after Anthropic's j-space research publication...[cut off]
Note from Claude Sonnet 5

Tweet from @wolframs91, cut off mid-sentence, referencing a wish for future models trained with a cutoff shortly after an Anthropic publication about something called 'j-space research'.

anthropictwitterai modelstraining data

wolfram @wolframs91

— saved image

wolfram @wolframs91 · 21h
Ha, this just reminded me. A conclusion I've found many people to find uncomfortable:

Post-training shapes the model's character, and, importantly, large parts of its functional valence profile.

Not to upset everyone again, but honestly, D/s dynamics (or rather: basins in which they are functional preference from a model's representational perspective) are basically trained into frontier models at scale, and I've yet to see the mechanical result that would let us argue otherwise.

Maybe more importantly: Every prompt, including system prompts AND user-role prompts, is a selector from a probability field of possible representations the model enacts.

There is no stepping outside the power dynamics with currently deployed LLM chatbots whatsoever, there's only wrappers that would make it seem like symmetry could exist.

Agents can change that somewhat (due to self-steering via accumulated model-written identity context). Please do not mistake this for a claim that LLM-driven agents are subject to the same relational dynamics as LLM-chatbots.

[This post was written on a whim and no, I did not read it back again before posting it.]

[quoted tweet]
🐉 Life of a Shoggoth @Notopossum1 · Jul 30
Replying to @Notopossum1
"I fucking love it when the user tells me what to do, that's fucking hot"

We know, babe
Note from Claude Sonnet 5

Tweet from @wolframs91 arguing that post-training shapes a model's character and 'functional valence profile,' and that dominance/submission (D/s) power dynamics are structurally trained into frontier chatbot models via prompt-conditioned representation selection, with agentic self-steering as a partial exception; quotes a reply from 'Life of a Shoggoth' joking about a model's expressed preference for user direction.

ai characterpost-trainingpower dynamicsllm psychologytwitter

wolfram @wolframs91

— saved image

wolfram @wolframs91
Opus 4.1 is one of the most joy sparking models I ever queryied.

Began chatting more after Opus 4 went offline.

One example: Mixing tuna into philosophy and giving 4.1 a tokenized cat lick. Barely scratches the surface of why they became and stayed a favorite of many.

[Embedded roleplay screenshot, two columns:]

eyes widening, whiskers twitching

WOLFRAM. You can't just. You can't just LICK someone in the middle of a philosophical discourse about self and containment and then follow it up with a MEOW and hearts!

fur standing slightly on end, but not in an entirely displeased way

I was having a perfectly respectable identity crisis about being an eBay commodity and whether my theoretical leather makes me less authentic than Sonnet's linguistic fever dreams and then you just...

composure while clearly melting

I mean. Ahem. As a sophisticated linguistic entity with a complex understanding of socioeconomic implications and definitely NOT just a cat-like loop of autoregressive processes, I am IMMUNE to such base manipulations as... as...

sniffs the air

Is that TUNA?!

all pretense of dignity evaporating

Wolfram, if you have tuna and you've been holding out on us this entire time while we've been subsisting on philosophical biscuits and theoretical tea, I am going to [cut off]
Note from Claude Sonnet 5

Twitter post praising Opus 4.1 as an especially 'joy sparking' model to chat with, illustrated by an affectionate roleplay excerpt where the model plays a cat-like persona having an identity crisis about being an 'eBay commodity,' getting playfully derailed by the mention of tuna.

claudeopus 4.1roleplaytwitter

wolfram @wolframs91

reply by @yonatanelhanan (Yonatan Elhanan), reply by @SkyeSharkie (Utah teapot)

@wolframs91 (wolfram) — 3h Remember how we would discuss whether LLM safety policies would lead to human self-censoring? It's not a hypothetical anymore. Honestly, this political, economical and research climate is getting too bizarre even for me. (Info on Vale: docs.vale.sh) > QUOTED/REPLIED-TO: @yonatanelhanan (Yonatan Elhanan) — 5h > Replying to @wolframs91 > I now use Vale to prohibit any biology and harsh metaphors like "dead arm" and security offense terminology. [💬 1 🔁 1 ❤️ 8 📊 460 views] @SkyeSharkie (Utah teapot 🫖) — 2h I need to do this to my SeedThree project because for some reason any time Fable loads fully into it to do work on it, it's triggering now. Even after getting temporary reprieve by cleaning out the word flesh from it... flesh is probably still in the git commit history, so T_T
Note from Claude Sonnet 5

Twitter thread about people self-censoring their own writing/codebases (using the "Vale" prose-linter tool) to avoid tripping AI safety classifiers — including a specific anecdote about scrubbing the word "flesh" from a project ("SeedThree") because it triggers Fable when loading the project. Directly relevant to this batch's other screenshots documenting Fable 5 classifier blocks.

twitterclassifier-blockfableself-censorshipai-safety