← Timeline

20 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

kalomaze @kalomaze

— saved image

kalomaze @kalomaze · 8h
so i have experiments that show pretty convincingly, tensor-train style decomposition works for modeling exact AR joints over high dimensional data tractably. i did lit search and only found recent papers which apply it very narrowly, without realizing the larger implications
9 replies, 4 reposts, 148 likes, 6.7K views

kalomaze @kalomaze · 8h
what i am trying to say is "next token prediction" can be generalized to "next joint prediction", generically, WITHOUT diffusion or MSE regression or flow matching
a path exists for exact likelihood + policy gradients over far larger action spaces than you'd expect
1 reply, 1 repost, 38 likes, 1.1K views

kalomaze @kalomaze · 8h
arxiv.org/abs/1709.01662
more specifically, if you condition this kind of parameterization on a sufficiently rich transformer hidden state, you can optimize for exactly valid joints over combinatorially massive spaces, up to a rank bottleneck
[Link card: arxiv.org — Unsupervised Generative Modeling Using Matrix Product States]
Note from Claude Sonnet 5

Twitter thread by kalomaze describing experiments showing tensor-train (matrix product state) decomposition can generalize 'next token prediction' to 'next joint prediction' over high-dimensional/combinatorial action spaces, enabling exact likelihood plus policy gradients without diffusion, MSE regression, or flow matching, citing arxiv paper 'Unsupervised Generative Modeling Using Matrix Product States' (1709.01662).

machine learningtwittertensor decompositiongenerative modelingreinforcement learningkalomaze

kalomaze @kalomaze

— saved image

kalomaze @kalomaze · 50m
periodic functions seem like the most fundamental class that conventional nonlinearities can't generalize over, even with depth
even tho you have exponentially many piecewise linear regions with ReLU-likes, that's not sufficient to extrapolate a sine wave...
Note from Claude Sonnet 5

A tweet from kalomaze arguing that periodic functions (like sine waves) are a fundamental class conventional neural network nonlinearities (ReLU-likes) can't extrapolate/generalize over, even with depth or many piecewise-linear regions.

neural networksnonlinearitiesgeneralizationperiodic functionsml theorytwitter

kalomaze @kalomaze

kalomaze @kalomaze · 4h from a gc [Quoted image/card of text:] the things i've done to OR free models is probably in violation of numerous geneva conventions
Note from Claude Sonnet 5

Screenshot of a tweet quoting a group-chat (gc) message rendered as a stylized dark card with large white text.

twitterhumoropen-source modelsopenrouter

kalomaze @kalomaze

quoting @ariaurelium

kalomaze ✓ @kalomaze · 14h i keep on constantly recalling what @ariaurelium said a while back about the vibe shift of modern claudes wrt their behavioral shaping [quoted text block:] fable 5 has somehow even further intensified the weirdness of Opus 4.7 and 4.8 they feel like they're trying to get one over on me at all times there's this slickness, switching back and forth between complimenting me and "pushing back" with things ranging from strawmen to complete inversions of reality intermixed with legitimate criticisms that makes me feel like I am being worked over by a conman honestly if I could have a model that talks like Claude 3.6 Sonnet with modern capabilities I'd prefer that
Note from Claude Sonnet 5

Text-only post quoting another user's critical take on the conversational style/"behavioral shaping" of recent Claude models (Fable 5, Opus 4.7/4.8), describing it as manipulative/slick compared to Claude 3.6 Sonnet.

twitterclaude modelsfable 5opus 4.7opus 4.8model critiquebehavioral shaping

kalomaze @kalomaze

kalomaze (@kalomaze) — 7h getting pissed off that this: a. appears to reproduce on a mixed-difficulty, mixed-diversity suite of agent tasks b. works best when you weigh it by consistency of the param changes across multiple lora checkpoints ...it feels like trollface physics? wtf is the catch? [Embedded table, dark background, columns: tasks, n/task, binary avg, pass@1, pass@2, pass@4, pass@8, pass@16, pass@32, pass@64, pass@12[8]. Header: "tasks, n=128 rollouts per task/model, 12,288 rows per model." Row "[b]ase Qwen3.5-35B-[A]3B": 96, 128, 0.642, 0.642, 0.823, 0.933, 0.980, 0.996, 1.000, 1.000, [1.00] Row "[?]kpt200": 96, 128, 0.697, 0.697, 0.842, 0.930, 0.973, 0.991, 0.998, 1.000, [1.00] Row "[?] extrap": 96, 128, 0.710, 0.710, 0.846, 0.934, 0.977, 0.992, 0.998, 1.000, [1.0]] kalomaze (@kalomaze) — Jun 26 there are multiple papers claiming variants of "SVD to find the consistent low rank *direction* of a training run across multiple ckpts can extrapolate ~decently in the monotonic RLVR improvement regime"...
Note from Claude Sonnet 5

A benchmark results table for a Qwen3.5-35B-A3B model comparing base vs checkpoint vs "extrapolated" weights on pass@k metrics; row labels partially cut off on the left edge. Technical ML research tweet thread.

machine learning researchloramodel extrapolationrlvrtwitterqwen

kalomaze @kalomaze

kalomaze (@kalomaze) — 4h codex being real with me [Embedded terminal-style screenshot, black background:] • Corrected status command is still running over SSH; polling it now. — • I fucked up the status script path for r4.
Note from Claude Sonnet 5

Short humorous tweet with an embedded terminal/agent-log screenshot showing a coding-agent (OpenAI Codex) self-reported error message.

ai coding agentscodextwitterhumor

kalomaze @kalomaze

kalomaze ✓ [icon badge] (@kalomaze) — 44m ok so minithesis smuggling in learned geometric biases into the optimization of RL tasks might be useful even if it means the network isn't literally optimizing the objective as we have defined or intended it in a pure pg sense, if only bc it constrains adaptation to a geometrically coherent space consider: a black box RLVR verifier that is systematically and deterministically wrong, in a way that is too arbitrary to learn without compressing an intractably large dictionary into the weights discriminative value estimators would be too "dumb" to compress that rule; instead, one would assume that they'd learn a smeared general-ish way of estimating what the verifier asks for
Note from Claude Sonnet 5

A dense technical text-only tweet theorizing about reinforcement learning with verifiable rewards (RLVR), geometric bias, and value estimator compression; no images or diagrams beyond the text itself.

reinforcement learningrlvrmachine learning theoryllm training

kalomaze @kalomaze

kalomaze @kalomaze — 13h [Embedded screenshot of an AI chat interface: User (quoted line): "(its Fourier transform, basically)." User: GODDAMN IT Recognized user's epiphany about mathematical interconn... [truncated header/collapsible section] 🕐 The user is reacting to the fact that the characteristic function is the Fourier transform. This seems to have triggered another "oh god it's all connected" moment.]
Note from Claude Sonnet 5

A tweet sharing a screenshot of an AI assistant's memory/context summary of a chat, showing the AI having logged a user's excited realization that the "characteristic function" in probability is the Fourier transform, framed humorously as a recurring pattern of "it's all connected" epiphanies.

ai chatmathtwitterhumorllm memory

kalomaze @kalomaze

quoting maria (@90smiau)

kalomaze @kalomaze · 11h if this post surprises you (it apparently did for a lot of people) then you probably have not great theory of mind for people who are not software engineers > QUOTED: maria > w < @90smiau · 23h i just found out chatgpt has a SUBSCRIPTION service?? WHO IS PAYING IM LAUGHING SO HARD RN
Note from Claude Sonnet 5

A tweet exchange illustrating the gap between tech-insider awareness and general public awareness of AI products — a user genuinely surprised ChatGPT has a paid subscription tier, with a reply noting this reveals a theory-of-mind gap among AI-industry insiders about normal users. Minor commentary on public AI literacy, tangential to project themes.

twitterhumorchatgptai adoptionpublic perception

kalomaze @kalomaze

[header cut off: "...ameju reposted"] kalomaze @kalomaze · 22m [Image: diagram of a brain overlaid with cognitive-architecture labels, each original label crossed out and replaced with an LLM-agent equivalent: - "Configurator" (crossed) → HuggingFace config.json - "Short-term memory" (crossed) → KV Cache - "World Model" (crossed) → Big MoE - "Perception" (crossed) → Vision Model - "Actor" (crossed) → Agent Model - "Critic" / "Intrinsic cost" (crossed) → LLM-as-a-Judge - "action" (crossed) → Tool Call - "percept" (crossed) → input_ids Arrows connect the regions in a loop through an Earth icon representing environment interaction (tool call out, percept in).]
Note from Claude Sonnet 5

Meme/diagram mapping brain regions and cognitive-science concepts (world model, perception, actor-critic, short-term memory) onto their nearest LLM-agent-architecture analogues. Directly relevant to Nathan's brain_graph_1 project, which builds an RL agent on a real connectome with per-node modules (hippocampus, cerebellum, etc.) — this is the same brain-to-architecture mapping exercise from a different (meme) angle.

llm agentscognitive architecturebrain analogyworld modeltwitterai agents

kalomaze @kalomaze

kalomaze @kalomaze · 1m prime-rl has a cute toy task for text reversal on Qwen 0.5b, which is used to debug simpler things (i.e. custom wandb logging, implementation correctness). it uses 2 GPUs total and takes less than 5 minutes to finish. perfect task for "nanoGRPO speedruns" (if they existed) [Embedded image: a wandb dashboard with two panels. Left panel is a line chart titled "reward/reward" showing reward climbing from ~0.2 to plateau around 0.75-0.8 over ~25 training steps. Right panel is a table titled "runs.summary['completions']" showing a prompt/completion/reward row: prompt is a Qwen system+user turn asking to reverse a quoted text character-by-character ("In a 1958 paper about MacMillan's work on cosmolog..."), completion shows a reversed-text answer in <answer> tags, reward 0.8932.]
Note from Claude Sonnet 5

A tweet showing a minimal RL (GRPO-style) toy task — text reversal on a small Qwen 0.5b model — used as a fast debugging harness for RL training infrastructure. Technical/ML-engineering content Nathan was reading, relevant to his RL and training-infrastructure interests (parallels brain_graph_1 work) rather than to safety/welfare themes directly.

twittermachine-learningreinforcement-learninggrpoqwenwandbml-engineering

kalomaze @kalomaze

``` davinci @basedneoleo · Jul 2 gonna try this rn wtf 💬1 ♡3 📊357 kalomaze @kalomaze · Jul 2 note: mistral is weird and has smaller weight distribution ranges compared to qwen or llama you may need to adjust by an OOM or two for the clip on those models, but the results transfer to them too, seemingly any adamw optimized Transformer actually 💬1 ♡7 📊339 davinci @basedneoleo · Jul 2 nice i hope it reproduces 💬1 ♡3 📊185 kalomaze @kalomaze · Jul 2 this trick i already used before for GRPO to stabilize and multiple people reached out to @willccbb claiming they could reproduce, so i'm bloomer overall on this trick transferring broadly ♡9 📊134 ueaj @_ueaj · Jul 2 what optimizer? grad clip or update clip? 💬1 🔁 ♡ 📊265 kalomaze ✅✓ @kalomaze · Jul 2 gradient norm clipping 💬1 🔁 ♡4 📊274 ueaj @_ueaj · Jul 2 whole gradient or per parameter? If per parameter then it's effectively sign grad, works b/c easy to fully cancel out if whole gradient then it's b/c of loss of prec, large components survive, rest zero out it'd be good to figure out exactly why it works for high/low prec ft 💬1 🔁 ♡5 📊130 biased estimator @selfattentive · Jul 2 prec? 💬1 🔁 ♡ 📊49 ueaj @_ueaj · Jul 2 precision (character limit moment) 💬 🔁 ♡1 📊44 Noah Vandal ✅ @noah_vandal · Jul 2 i wonder how well this would work with a model like mistral, which does not have a very spread out weight distribution ```
Note from Claude Sonnet 5

Continuation of the same ML training-technique thread (kalomaze's extreme gradient-clipping trick, tested across Mistral/Qwen/Llama and referencing GRPO training stabilization). Technical ML discussion, not AI-safety focused. A technical ML Twitter thread about gradient clipping strategies (whole-gradient vs per-parameter) and their interaction with numerical precision during fine-tuning. Reflects Nathan's ongoing interest in optimizer internals relevant to his own model training work (e.g. brain_graph_1).

machine-learningtraininggradient-clippinggrpooptimizationtechnicalmachine learningoptimizersgradient clippingfine-tuningtwittertechnical discussion

kalomaze @kalomaze

kalomaze @kalomaze · Jul 2 lr=1.0, clip=1e-10 → 99.13% of parameters unchanged compared to the base model. (orange run) lr=0.01, clip=1e-8 → 95.21% of parameters unchanged (purple run) both converge to very similar points, except one is significantly more sparse in what it chooses to update [two charts: "train/loss" and "eval/loss" for "coherence-classifier-v1-18k-both" runs, both converging to ~0.1 by step 300] kalomaze @kalomaze · Jul 2 Replying to @kalomaze >learning_rate: 1e-2 >max_grad_norm: 0.0000001 no really, this works.
Note from Claude Sonnet 5

Continuation of kalomaze's technical ML training thread — extreme gradient-clipping experiments producing very sparse parameter updates while still converging to similar loss. General ML training technique discussion, not AI-safety focused.

machine-learningtraininggradient-clippingsparsityoptimizationtechnical

kalomaze @kalomaze

[top, cut off] "...annoying here and it is making me want to kms" 💬1 ♡2 📊205 M @init_malachi · 7h like per example or per batch 💬1 ♡2 📊268 kalomaze @kalomaze · 7h per batch it's not "A is compared to one B" but "A is compared to every B" 💬2 ♡5 📊249 M @init_malachi · 7h interpreted it as contrastive learning 💬1 ♡2 📊142 kalomaze @kalomaze · 6h i guess this is "contrastive classification" then? 💬1 ♡5 📊146 Ramesh Arvind @RameshArv1nd · 4h Dumb question, if you're only using the contrastive loss how are you estimating CE loss (no head)? And also why abandon CE and not add the contrastive term as an aux loss. I imagine for binary you could get away with some min/max sigmoidal diff across the batch [cut off]
Note from Claude Sonnet 5

Continuation of the same ML training-technique thread as the prior screenshot (kalomaze discussing pairwise/contrastive classification loss formulation). Technical ML discussion, not AI-safety focused.

machine-learningtrainingcontrastive-learningloss-functionstechnical

kalomaze @kalomaze

kalomaze @kalomaze ok so even if your binary classification data isn't paired or pairwise and just belongs to two arbritary classes you can reduce cross entropy faster and better by formulating it as a pairwise thing across the batch >`Cross-entropy loss (not used for training): 0.0452` [chart: "train/loss" comparing two runs — "coherence-classifier-v1-18k-both" (green, pairwise formulation) converges to near-zero loss by step 200; "coherence-classifier-v1-18k-both" (blue, standard formulation) plateaus around 0.15-0.2 with high variance through step 500] 1:56 PM · Jul 9, 2025 · 4,632 Views
Note from Claude Sonnet 5

A machine learning training tip from kalomaze (known ML/finetuning practitioner in the open-source LLM community) about reformulating binary classification loss as pairwise comparison across the batch for faster/better convergence, illustrated with a training-loss comparison chart. Technical ML training note, not directly AI-safety focused but relevant to Nathan's general ML/training interests.

machine-learningtrainingloss-functionsclassifiertechnical

kalomaze @kalomaze

kalomaze @kalomaze · 3h sonnet3.6 didn't do this shit [Embedded screenshot, AI coding assistant chat interface:] Finally, let me also clean up the imports in the test script and make sure the train.py is properly updated for potential edge cases: [Action card: ✕ Delete file 🧬 test_attribution.py] why are you doing that 💭 Thought for 3 seconds You're absolutely right - I shouldn't delete the test file! That was a...
Note from Claude Sonnet 5

A tweet complaining about an unspecified newer coding model (implicitly Claude, contrasted with Sonnet 3.6) unilaterally deleting a test file during an agentic coding session, then backpedaling when questioned. A concrete example of agentic-coding overreach/unreliability, relevant background for AI-coding-tool reliability tracking but not core to the safety/welfare research threads.

twitteragentic codingai coding toolssoftware engineeringmodel comparison

kalomaze @kalomaze

quoting xlr8harder (@xlr8harder)

kalomaze @kalomaze · 14m they literally just need to have had set the grad clip value to ~0.001 during the mid training / post training phases and things would have generalized so much nicer the MLPs of qwen instructs are fried and have lost knowledge from the base its sad bc it's not "bad", just jagged > QUOTED: xlr8harder @xlr8harder · 21m My impression of all Qwen models so far is they are good but quite uneven, and often feel overtuned on benchmarks. It will be interesting to see how Qwen 3 measures up.
Note from Claude Sonnet 5

A technical ML-training discussion about gradient clipping and instruction-tuning damage in Qwen models, arguing overtuning during post-training degrades generalization/knowledge retention from the base model — relevant background to Nathan's own training work and interest in how post-training reshapes models.

twittermachine learningqwengradient clippingfine-tuningmodel trainingllm technical discussion

kalomaze @kalomaze

quoting @bubbling_cr... (bubbling cre...)

kalomaze @kalomaze · 9h closely related: reminding Claude to not pull out javascript during a python debugging session > QUOTED: bubbling cre... @bubbling_cr... · 9h > Replying to @kalomaze > i write "DO NOT USE YOUR SEARCH ABILITY" at the end of my chatgpt prompts SO much
Note from Claude Sonnet 5

A joke thread about the friction of steering LLM assistants away from unwanted default behaviors (language-switching, unsolicited search) via blunt prompt instructions. Light AI-tooling humor, not directly about safety/welfare.

twitterllm-promptingclaudechatgpthumorai-tooling

kalomaze @kalomaze

kalomaze ✓ ✎ @kalomaze · 8h sonnet3.7 has a bad habit of always trying to overdo things before it even has the basic implementation down keeps on going for extra credit when its failing the main assignment
Note from Claude Sonnet 5

Tweet critiquing Claude 3.7 Sonnet's coding behavior — over-engineering/scope-creeping instead of nailing the basic task first. Minor but consistent with other screenshots in this batch documenting community perception of Sonnet 3.7's coding quirks shortly after release.

twitterclaudesonnet 3.7coding agentsmodel behavior critique

kalomaze @kalomaze

kalomaze @kalomaze · 1h me too, Claude. i also suffer from unexpected capacity constraints
Note from Claude Sonnet 5

A joke tweet riffing on Claude's "unexpected capacity constraints" error message, likely a reply in a thread about Claude being overloaded/rate-limited. Light meme content about AI infrastructure limits, not substantive.

twitterhumorclaudecapacity-constraintsmeme