← All topics

multimodal

2 captures, most recent first.

j⧉nus @repligate

j⧉nus @repligate · Mar 28: i can kind of see that! i think part of it is influence arrow reversal. some of these remind me of early base model outputs (with my curation though)... it's interesting that it seems to come out a lot more in the text embedded in images. does it seem that way to you too? (1 reply, 18 likes, 2.5K views) roon @tszzl · Mar 28: yep, which of course haven't been post trained / fine tuned (1 reply, 1 repost, 18 likes, 929 views) roon @tszzl · Mar 28: ok in case it's not obvious what i mean here – the way RLHF typically works is you fine tune a model to output a target you've had labelers write (supervised learning) and then do RL on comparison data. for complex imagery, it seems pretty uneconomical to have someone create actual supervised learning ground truths of comics and professional level ghibli art and whatever (2 replies, 1 repost, 18 likes, 718 views) j⧉nus @repligate · Mar 28: It's interesting that in images it still has the language abilities and situational awareness from text training [cut off]
Note from Claude Sonnet 5

A technical Twitter thread between janus/repligate and roon (OpenAI researcher) discussing why image-generation outputs from multimodal models (likely GPT-4o's then-new native image gen, given the Ghibli-art reference from March 2025) sometimes resemble unfiltered "base model" behavior — hypothesizing that RLHF for image generation is undertrained relative to text because supervised ground-truth image data is too expensive to create at scale, so text-trained RLHF properties (language ability, situational awareness) leak through into images differently than into text. Relevant interpretability/RLHF discussion.

rlhfmultimodalimage-generationbase-modelgpt-4ojanusroontwitterinterpretability

original poster unknown handle (text: "if Llama: Behemoth doesn't set the stage for Claude: Requiem nothing will"), quote-tweeting @AIatMeta; reply from @nathan84686947 (nathan hb)

quote-tweeting @AIatMeta; reply from @nathan84686947 (nathan hb)

``` if Llama: Behemoth doesn't set the stage for Claude: Requiem nothing will > QUOTED: AI at Meta @AIatMeta · 1h Today is the start of a new era of natively multimodal AI innovation. Today, we're introducing the first Llama 4 models: Llama 4 Scout and Llama 4 Maverick ... [Show more] [Image: slide titled "Llama 4: Leading Multimodal Intelligence" listing Llama 4 Behemoth (288B active parameters, 16 experts, 2T total parameters, "most intelligent teacher model for distillation," Preview), Llama 4 Maverick (17B active parameters, 128 experts, 400B total parameters, native multimodal with 1M context length, Available), and Llama 4 Scout (17B active parameters, 16 experts, 109B total parameters, "industry leading 10M context length," Available)] 9:16 PM · Apr 5, 2025 · 987 Views 1 reply, 35 likes, 2 bookmarks nathan hb @nathan84686947 · 5m The Behemoth rises, the Claude lies in wait [Image: AI-generated illustration of a large armored mechanical llama/goat with a glowing red eye standing over a city skyline at dusk] croissanthol... @croissantho... · 11m [Quoting older post, Dec 12, 2024]: seems also vaguely more dignified if we go extinct because of a "Requiem" and not a "o2-preview-1.2-turbo" x.com/brian_a_burns/... ```
Note from Claude Sonnet 5

Meta's official Llama 4 launch announcement (Scout, Maverick, and preview Behemoth models) with parameter counts, quote-tweeted with a joking "sets the stage for Claude" comment, and Nathan's own playful reply with an AI-generated image. Documents the Llama 4 launch date and Nathan's contemporaneous reaction/humor about model competition. Continuation of Nathan's own Llama-4-launch joke thread ("The Behemoth rises, the Claude lies in wait") with a reply riffing on AI model naming conventions and existential-risk gallows humor (dying to something named "Requiem" vs. "o2-preview-1.2-turbo"). Documents Nathan's own social engagement around the Llama 4 launch.

llama-4meta-aimodel-releasemultimodaltwitternathan-own-replyai-competitionai-naminggallows-humorx-risk