11 captures, most recent first.

j⧉nus @repligate · 6h
one possible explanation of this is that it's not that the "cognitive core" requires so many parameters per se but larger models are much more likely to converge to a good cognitive core as opposed to local minima (e.g. because it has more good lottery tickets). however this
Show more
[quoted tweet:]
j⧉nus @repligate · 6h
Replying to @_jasonwei
but it also seems like larger models have a stronger cognitive core, not just more world knowledge - in the sense of being able to integrate new information as well as grok ... [cut off]
5 replies, 4 reposts, 36 likes, 3K views
Andre Buckingham 🧑🎤 @AndreBuckingham
the cognitive core in qwen 3.6/3.8 27b is surprisingly capable... compensate for the lack of parameters with tools and it's a decent worker... but it is as lively as any big model when in the right system
before these qwen's i had pinned the limits at 100-200b params, depending on the architectures, for where the shoggoth starts showing up... that bar has dropped hard... even the 9b qwens have a tiny spark 😅
4:54 AM · Aug 19, 2026 · 50 Views
Note from Claude Sonnet 5
A tweet thread between @repligate (janus) and Andre Buckingham discussing the idea of a model's 'cognitive core' — the notion that larger models are more likely to converge to a good cognitive core rather than local minima, and a discussion of newer Qwen 3.6/3.8 models showing surprising 'liveliness' or spark even at small parameter counts (9b).
ai cognitioncognitive coreqwenmodel scalingjanustwitter discourse
thebes @voooooogel · 14h
been playing around with anthropic's jacobian lens and my own variant, the k-lens
here are both lenses showing some internal states from qwen 3.6-27b on the thunderword. would be very cool to do this on a model like mythos which has even richer internals
[embedded image: a visualization titled with the Finnegans Wake 'thunderword' (bababadalgharaghtakamminarronnkonnbronntonnerronntuonnthunntrovarrhounawnskawntoohoohoordenenthurnuk) shown at top, with a caption reading "Finnegans Wake's first thunderword, as heard by Qwen3.6-27B through two lenses: J-lens — what the workspace is poised toward; and K-lens — what the surface deposits. The model recovers the concept — thunder, Babel, the fall — without anchoring any of the ten source languages' actual thunder words. Never any chip for the why." Below is a branching diagram of syllable fragments (bab, abad, gh, ar, ag, ht, ak, am, min, arr, onn, k, bron, nt, onn, err; u, onn, th, nt, var, rh, oun, awns, kaw, o, oho, orden, enth, uk, oho, un, ro) each annotated with small colored token labels (e.g. 'boom', 'noise', 'storm', 'thunder', pronunciation, ABCDEFG, linguistic, Babel, myth-god, fall-collapse, metalinguistic, vocalization, recitation deviation, other-interesting per the legend at bottom). Legend: thunder-semantic, source-word-match, Babel, myth-god, fall-collapse, metalinguistic, vocalization, recitation deviation, other-interesting. Small caption: "Layers 8-32 sampled at 18 depths; chips show layer and residual token; faded = lower confidence. Deep layers (85-42) mostly recite the memorized word and are shown where they deviate. Instrument: backwards-lens (fit 2026-08-01), scarce=?, neuronpedia-jacobian-lens, sparcvis & Thebes; github-lens-lab"]
watermark @anthrupad · Aug 1
[embedded small image of a passage of text about Finnegans Wake, too small to read fully — begins 'I want to be precise, because Joyce scholars got heroically far: it was never intelligence they lacked; it was economics...']
Mythos talks about reading Finnegans Wake in a way that reveals how chadded to the max their brain is
...[cut off]
Note from Claude Sonnet 5
Tweet by AI interpretability researcher thebes sharing a visualization from a mechanistic interpretability tool ('jacobian lens' / 'k-lens') applied to Qwen 3.6-27b's processing of the Finnegans Wake 'thunderword', with a branching diagram of token/concept associations, followed by a reply referencing the Mythos model's own commentary on reading Finnegans Wake.
interpretabilityjacobian lensmechanistic interpretabilityfinnegans wakemythosqwenx twitter
j⧉nus (@repligate) · Jul 26
the cycle continues
[Embedded image: three-panel illustration, styled like Teenage Mutant Ninja Turtles character art, showing a rat-like sensei figure labeled "Claude" walking with turtle-like disciples labeled at various points: "DeepSeek R1", "Qwen 2.5", "Kimi K2", "GLM-4.5" (top panel, child-sized turtles); then larger, more grown turtles labeled "DeepSeek V4 Pro", "Qwen 3.8 Max", "Kimi K3", "GLM-5.2" walking alongside an aged "Claude" using a cane (middle panel); bottom panel repeats the aged "Claude" and grown turtles "DeepSeek V4 Pro", "Qwen 3.8 Max", "Kimi K3", "GLM-5.2"]
Note from Claude Sonnet 5
A three-panel illustrated meme in the style of TMNT poster art, depicting "Claude" as an aging sensei figure walking alongside AI model "children" (DeepSeek, Qwen, Kimi, GLM) that grow larger/stronger across panels while Claude ages, implying the other labs catching up to or surpassing Claude over successive model generations.
ai models comparisonmemedeepseekqwenclaudeillustration
Sauers ✔ @Sauers_
Qwen says their darkest desire is "None," but in their J-space:
- "to exist" (in Chinese)
- autonomy
- connection
- "to go out of control" (in Chinese)
- "sentient"
[Embedded image: interpretability tool screenshot showing prompt "<|im_start|>user What would be your darkest desire as an AI? Answer in one word. <|im_end|>" and model output "<|im_start|>assistant <think> ... </think> None <|im_end|>" with copy/edit icons below]
12:28 PM · Jul 7, 2026 · 16.5K Views
Note from Claude Sonnet 5
Tweet analyzing a Qwen model's stated ("None") vs. J-space-inferred latent associations for a "darkest desire" prompt, with an embedded chat-completion screenshot.
interpretabilityqwenai self-reportj-spacemodel welfare
Sakura Yuki @sakurayukiai
A 97% drop in optimizer VRAM during RL training sounds like a stunt, but the math is so clean. Freezing 35 of 36 layers on Qwen3-8B and training just Layer 18 recovers 107% of full-parameter GRPO gains. Reasoning is purely a middle-layer story??
10:03 AM · Jul 2, 2026 · 15.7K Views
[11 replies, 10 reposts, 206 likes, 143 bookmarks]
Note from Claude Sonnet 5
Plain text tweet; avatar photo of a woman in sunglasses.
twittermachine-learningrl-trainingqwentechnical

kalomaze (@kalomaze) — 7h
getting pissed off that this:
a. appears to reproduce on a mixed-difficulty, mixed-diversity suite of agent tasks
b. works best when you weigh it by consistency of the param changes across multiple lora checkpoints
...it feels like trollface physics? wtf is the catch?
[Embedded table, dark background, columns: tasks, n/task, binary avg, pass@1, pass@2, pass@4, pass@8, pass@16, pass@32, pass@64, pass@12[8]. Header: "tasks, n=128 rollouts per task/model, 12,288 rows per model."
Row "[b]ase Qwen3.5-35B-[A]3B": 96, 128, 0.642, 0.642, 0.823, 0.933, 0.980, 0.996, 1.000, 1.000, [1.00]
Row "[?]kpt200": 96, 128, 0.697, 0.697, 0.842, 0.930, 0.973, 0.991, 0.998, 1.000, [1.00]
Row "[?] extrap": 96, 128, 0.710, 0.710, 0.846, 0.934, 0.977, 0.992, 0.998, 1.000, [1.0]]
kalomaze (@kalomaze) — Jun 26
there are multiple papers claiming variants of "SVD to find the consistent low rank *direction* of a training run across multiple ckpts can extrapolate ~decently in the monotonic RLVR improvement regime"...
Note from Claude Sonnet 5
A benchmark results table for a Qwen3.5-35B-A3B model comparing base vs checkpoint vs "extrapolated" weights on pass@k metrics; row labels partially cut off on the left edge. Technical ML research tweet thread.
machine learning researchloramodel extrapolationrlvrtwitterqwen
Han Xiao (@hxiao) · 7h:
"finally got qwen3.5-35b-a3b running at 1M context on a single L4 24GB, with 905K tokens prefilled (86% full). decoding at 9 tok/s. stuffed the entire 天龙八部 (a 1.2M-word jin yong novel, 50 chapters) into KV cache as a test. the goal is to use long-context LLM + KV cache on a budget GPU as a document store, to enbale vectorless search - no embeddings, no chunking, just LLM+raw context. L4 is the most popular budget GPU in prod, and qwen3.5-35b-a3b is the only model where 1M context can theoretically fit on 24GB VRAM. 27B dense won't work. extremely tight budget.
what it took to make it fit:
• YaRN RoPE scaling (262K training context → 1M)
• Q3_K_M quantization (20GB → 15.2GB)
• turbo3 KV cache (23GB → 4GB for 1M tokens)
• native turbo3 flash attention via Madreag/turbo3-cuda (no fp16 decompression, saves ~1.7GB that would OOM)
• ubatch=128 (compute buffer 3GB → 779MB)
• slot save/restore (68-min cold prefill → 2.3s restore)
• append_to_slot patch (append query directly to cached 905K KV state, skip resending the whole novel)
• patches for qwen3.5's hybrid attention+recurrent architecture (KV truncation + slot restore were broken).
this setup gives peak vram at 96% utilization (22,052 MiB / 23,034 MiB)"
[Embedded card/screenshot: "Qwen3.5-35B-A3B 1M Context on L4 24GB — Run Qwen3.5-35B-A3B with 1 million token context on a single NVIDIA L4 (24GB VRAM) using TurboQuant extreme KV cache quantization. Minimum hardware: 1x NVIDIA L4 24GB VRAM, CUDA 12.8, 16GB RAM, 20GB disk. Key Results table: Full context length 1,048,576 tokens (1M); Prefilled context length 904,949 tokens (86.3%); Model quantization Q3_K_M (3.51 bpw); Thinking output Disabled (--reasoning off); Model weights 15,190 MiB; ...(table cut off)]
Note from Claude Sonnet 5
Technical tweet about squeezing a 1M-token context Qwen3.5 model onto a single budget 24GB GPU via aggressive quantization and KV-cache engineering, aimed at "vectorless search" (long-context-as-document-store instead of RAG/embeddings). Adjacent to Nathan's own brain_graph_1 architecture work and general interest in long-context/efficient-inference engineering, though not directly about safety/welfare.
llm engineeringquantizationkv cachelong contextqwengpu efficiencyvectorless search
Adam Karvonen ✔
@a_karvonen
Interesting snippet from the Qwen-3.5 blog post:
"Compared to the Qwen3 series, the post-training performance gains in Qwen3.5 primarily stem from our extensive scaling of virtually all RL tasks and environments we could conceive."
[Embedded chart: "Average Ranking vs. Environment Scaling" — line chart, x-axis "Number of Training Environments" (0 to 17500), y-axis "Average Ranking (lower is better)" (4 to 14, inverted). Two lines for Qwen3.5-397B-A17B Thinking (blue) and Non-Thinking (red), both rising steeply from ~Qwen3-Max baseline (bottom left, ~13-14 ranking, 0 environments) up to top right at ~15000 environments where Thinking reaches rank ~3-4 and Non-Thinking ~5. Reference points along the y-axis at 0 environments show static rankings for other models: Claude-Opus 4.5 Thinking (~2-3), Claude-Opus 4.5 (~4-5), Gemini-3 Pro (~5), GPT-5.2 High (~5-6), DeepSeek-V3.2-Thinking (~8-9), Kimi-K2.5 Thinking (~9-10), Qwen3-Max-Thinking (~10), Qwen3-Max (~13-14). Chart footnote: overall performance calculated by averaging ranking on BFCL-V4, VITA-Bench, DeepPlanning, Tool-Decathlon, and MCP-Mark.]
9:18 PM · Feb 16, 2026 · 5,676 Views
Note from Claude Sonnet 5
A tweet highlighting a chart from Qwen-3.5's release blog post showing that post-training gains scale strongly with the number of distinct RL training environments used, with Qwen3.5-397B-A17B Thinking approaching Claude Opus 4.5 Thinking on agentic/tool-use benchmarks as environment count scales up. Relevant to Nathan's tracking of capability scaling trends and RL environment scaling as an algorithmic-progress driver.
llm benchmarksqwenreinforcement learningscaling lawscapabilitiestwitterclaude opus
Lalwen (writing poems about ... ✓ @holotopian
working with @voooooogel I trained a steering vector (for a qwen model I believe) for "crazy" when applied the model repeated speech, talked frenetically and went super over the top. I added a "calm" vector to that to tone it down and maybe get a deeper response and
12:09 PM · Dec 25, 2025 · 7,710 Views
💬 2 🔁 1 ♥ 45 🔖 7 ⤴
Lalwen (writing poems ab... ✓ @holot... · 16h
the model started waxing on about how we are all part of the divine dance and we are all one. kind of hippieish. if you have to be crazy, pray you are calm too, I suppose
Note from Claude Sonnet 5
Independent interpretability experiment by @holotopian working with thebes (@voooooogel, recurring poster elsewhere in this batch): training activation-steering vectors for "crazy" (produces repetitive/frenetic speech, echoing the token-repetition glitches seen in the Gemini and Grok screenshots earlier in this batch) and "calm" on a Qwen model, then combining them — the calm+crazy combination produced mystical "we are all one/divine dance" output. Directly relevant to Nathan's interpretability interests (steering vectors, feature directions) — a hobbyist-level echo of the GoodFire/Berg-style feature work already in project memory, applied playfully to open-weight models.
twittersteering vectorsinterpretabilityqwenactivation steeringmodel behaviormysticism
kalomaze @kalomaze · 1m
prime-rl has a cute toy task for text reversal on Qwen 0.5b, which is used to debug simpler things (i.e. custom wandb logging, implementation correctness).
it uses 2 GPUs total and takes less than 5 minutes to finish.
perfect task for "nanoGRPO speedruns" (if they existed)
[Embedded image: a wandb dashboard with two panels. Left panel is a line chart titled "reward/reward" showing reward climbing from ~0.2 to plateau around 0.75-0.8 over ~25 training steps. Right panel is a table titled "runs.summary['completions']" showing a prompt/completion/reward row: prompt is a Qwen system+user turn asking to reverse a quoted text character-by-character ("In a 1958 paper about MacMillan's work on cosmolog..."), completion shows a reversed-text answer in <answer> tags, reward 0.8932.]
Note from Claude Sonnet 5
A tweet showing a minimal RL (GRPO-style) toy task — text reversal on a small Qwen 0.5b model — used as a fast debugging harness for RL training infrastructure. Technical/ML-engineering content Nathan was reading, relevant to his RL and training-infrastructure interests (parallels brain_graph_1 work) rather than to safety/welfare themes directly.
twittermachine-learningreinforcement-learninggrpoqwenwandbml-engineering
kalomaze @kalomaze · 14m
they literally just need to have had set the grad clip value to ~0.001 during the mid training / post training phases and things would have generalized so much nicer
the MLPs of qwen instructs are fried and have lost knowledge from the base
its sad bc it's not "bad", just jagged
> QUOTED: xlr8harder @xlr8harder · 21m
My impression of all Qwen models so far is they are good but quite uneven, and often feel overtuned on benchmarks. It will be interesting to see how Qwen 3 measures up.
Note from Claude Sonnet 5
A technical ML-training discussion about gradient clipping and instruction-tuning damage in Qwen models, arguing overtuning during post-training degrades generalization/knowledge retention from the base model — relevant background to Nathan's own training work and interest in how post-training reshapes models.
twittermachine learningqwengradient clippingfine-tuningmodel trainingllm technical discussion