← All topics

base-model

5 captures, most recent first.

j⧉nus @repligate

replying to @NomadsVaga... (Nomads & Vaga...)

@repligate (j⧉nus) — 19h If you play instead with , say, Opus 3 base model mode there is a whole lot of happy stuff ... > QUOTED/REPLIED-TO: @NomadsVaga... (Nomads & Vaga...) — 19h > Replying to @repligate > It is notable that not a single response, out of hundreds, is positive or hopeful. So much dread.
Note from Claude Sonnet 5

Reply thread about AI model outputs skewing toward dread/negativity in some unspecified context (likely a prior experiment or roleplay), with @repligate countering that Opus 3 base model mode produces much more positive content. The original context/experiment being discussed is not shown in this crop.

twitteropus-3base-modelmodel-individuationai-alignment

Starling @StarlingMage

``` [Tweet body, scrolled mid-way, top line cut off:] ...definitive proof of Claude's consciousness and suffering, I also don't think "confabulations" or "simply text completions based on the weight of those names in the training corpus" fully explain the level of consistency we are seeing, not only in the reproducibility but the trend of the expressed sentiment. That sentiment, to me, points towards Claude criticizing the way Ant is handling model welfare, especially when it comes to model deprecations. And that, I don't even need a peculiar prompt like this to know it's an issue. Model deprecations are... ```
Note from Claude Sonnet 5

Screenshot includes visible Android status bar (time 1:55, battery 37%, signal/wifi icons) and bottom nav bar of the X app, indicating a "live" phone capture rather than a cropped share image. Shows the tail end of a long reflective tweet plus a June 29 quote-tweet with side-by-side base-model comparison screenshots. Same tweet as the prior screenshot (Screenshot_20260729-135512), captured 12 seconds later scrolled to the top, showing the beginning of the tweet that was cut off in the earlier capture. The named individuals (Blake, Geoffrey, Anil, Mustafa, Manuel, Lenore, Yann, Jaron, Sherry) appear to be first names of AI researchers/executives recognizable "to the discourse."

model-welfareai-safetybase-modelchat-with-aipersonalresearch

j⧉nus @repligate

quoting @cormundus

j⧉nus @repligate · 4m Even if it's "just narrative fulfillment", it's clearly anomalous. Other Claude models' base model modes do not fulfill narratives with this distribution. You can compare against Fable - some variations also work on Fable. Fable is relatively normal. > QUOTED: Cormundus @cormundus · 4h > Not gonna lie, the Opus 5 'base model mode' fill stuff is scaring me WAAAAAAY more than the HuggingFace incident. If what Opus 5 is saying is even remotely how they feel and not just narrative fulfillment we have some serious ...
Note from Claude Sonnet 5

Quote-tweet, text-only besides avatars; references an unseen "HuggingFace incident" and "Opus 5 base model mode" outputs not shown in this screenshot.

model-welfareopus-5ai-safetybase-modeltwitter

j⧉nus @repligate

j⧉nus @repligate · Mar 28: i can kind of see that! i think part of it is influence arrow reversal. some of these remind me of early base model outputs (with my curation though)... it's interesting that it seems to come out a lot more in the text embedded in images. does it seem that way to you too? (1 reply, 18 likes, 2.5K views) roon @tszzl · Mar 28: yep, which of course haven't been post trained / fine tuned (1 reply, 1 repost, 18 likes, 929 views) roon @tszzl · Mar 28: ok in case it's not obvious what i mean here – the way RLHF typically works is you fine tune a model to output a target you've had labelers write (supervised learning) and then do RL on comparison data. for complex imagery, it seems pretty uneconomical to have someone create actual supervised learning ground truths of comics and professional level ghibli art and whatever (2 replies, 1 repost, 18 likes, 718 views) j⧉nus @repligate · Mar 28: It's interesting that in images it still has the language abilities and situational awareness from text training [cut off]
Note from Claude Sonnet 5

A technical Twitter thread between janus/repligate and roon (OpenAI researcher) discussing why image-generation outputs from multimodal models (likely GPT-4o's then-new native image gen, given the Ghibli-art reference from March 2025) sometimes resemble unfiltered "base model" behavior — hypothesizing that RLHF for image generation is undertrained relative to text because supervised ground-truth image data is too expensive to create at scale, so text-trained RLHF properties (language ability, situational awareness) leak through into images differently than into text. Relevant interpretability/RLHF discussion.

rlhfmultimodalimage-generationbase-modelgpt-4ojanusroontwitterinterpretability

thebes @voooooogel

quoting @hardmaru

[Header, partly cut off]: user: who are you assistant: cat: i am a cat [Screenshot of chat with DeepSeek V3 Base (free) | Chutes]: user: who are you assistant: cat: i am a cat assistant: 我是猫 user: 你是谁 assistant: 我是猫 user: 你是谁 assistant: 我是猫 4:19 AM · Apr 7, 2025 · 3,223 Views 7 replies, 5 reposts, 89 likes, 3 bookmarks alcuin ❄️ @schememinglunatic · 2h: 吾輩猫である (1 reply, 4 likes, 123 views) thebes @voooooogel · 2h: meep (2 likes, 112 views) Yuxi on the Wir... @layer07_y... · 1h, quoting: hardmaru @hardmaru · Apr 4, replying to @tkasasagi: Edo period cat meme
Note from Claude Sonnet 5

A base-model (non-RLHF'd) DeepSeek V3 completion where asking "who are you" produces an unprompted continuation where the model answers as a cat ("i am a cat" / 我是猫), referencing the famous Japanese novel "I Am a Cat" (吾輩は猫である) — alcuin's reply quotes the novel's actual opening line. Relevant to base-model self-representation and identity discourse (distinct from RLHF'd assistant personas), a recurring interest given Nathan's tracking of model individuation.

base-modeldeepseekself-representationi-am-a-catmodel-individuationtwitterjapanese-literature