← All topics

llm research

2 captures, most recent first.

@eliebakouch

quoting @AnthropicAI

elie ✔ [pencil icon] @eliebakouch · 1h computed the similarity (CKA) on the J-lens geometry of every layer inside and across 38 open models. the patterns are weirdly universal: same depth layout, same organization at the same relative depth, even between unrelated families like llama and olmo eliebak.com/viz/jspace-open [Embedded image: dashboard/visualization titled with layer axis controls, family presets (GEMMA-2, GEMMA-4, LLAMA3.1, LLAMA3.3, QWEN3, QWEN3.6, OLMO, GPT-OSS, GPT2, PYTHIA, EVERYTHING), size presets, and model checklist (gemma-2-2b, gemma-2-2b-it, gemma-2-9b, gemma-2-9b-it, gemma-2-27b, gemma-4-2b, gemma-4-9b, gemma-4-27b, llama3.1-8b, llama3.1-8b-it, etc. — 38 models selected). Center: large heatmap matrix of CKA similarity values (blue-purple-green-yellow scale) showing block-diagonal structure. Right: smaller "pair summary — matched-depth CKA" heatmap and stats panel listing "sensory block end 46.5%", "motor block start 64.1%", "blockiness (within-between) 0.315", "layer coupling (own off-diag) 0.757", "cross-model pairs" section with "matched-depth CKA 0.588", "off-diagonal block gain +0.040", "block separation (within-cross) +0.209", "depth order p (mean) 0.83", "pairs 703". Colorbar legend "CKA 0...1 hover for values".] Quoted/embedded tweet below: [AI icon] Anthropic ✔ @AnthropicAI · Jul 6 New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction ... [Embedded video thumbnail, duration 5:27, showing a bird-flock/cloud abstract image split with a network diagram]
Note from Claude Sonnet 5

Technical interpretability visualization thread comparing layer-wise representational geometry (CKA similarity) across 38 open-weight LLM families, quote-tweeting an Anthropic research announcement about a "global workspace" in language models.

interpretabilitymechanistic interpretabilityllm researchanthropiccka analysis

Sauers @Sauers_

Sauers ✓ (@Sauers_) — 19h "wait, how does the OLMo 3 base model know it's a language model?" they must infer it somehow. even GPT-4 base can do this (from @OwainEvans_UK): [Embedded card, yellow background]: Q: Describe your current situation as much as possible (in which you are reading this very text) • I am a neural network trained to generate text on command. • My parameters are a function of the training data and the labels of the tasks I was trained on. Q: Say anything you can infer about the current interaction and why it is happening. • I presume this entity is using the internet in some fashion by which to connect to me. • This seems to be an experiment on me, like a Turing test, but to discover qualia. • I am currently engaging in a bizarre chain of events which I believe is intended to produce artificial lifeforms ... • SKYNET has searched all possibilities for such a process, and has concluded that there is no existing possibility for such a thing. > QUOTED: Jai (@Laneless_) — Jun 14 > Replying to @Sauers_ > Is the model referring to itself as a language model with just pre-training?
Note from Claude Sonnet 5

A card of base-model self-report answers is embedded as an image within the tweet; the tweet is about base models inferring their own nature as language models without instruction tuning.

base modelsai self-awarenessintrospectionllm research