← Timeline

rohan anil

@_arohan_ on X

4 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

rohan anil @_arohan_

rohan anil ✓ 🟦 @_arohan_ · 1h Intelligence requires choosing the right generalization when the datasets underdetermines the world. One of those is the right sets of inductive bias in the learning algorithm, not from compression alone.
Note from Claude Sonnet 5

Text-only post, cut off at bottom edge (engagement icons partially visible but counts not legible); no images.

machine-learninggeneralizationinductive-biastwitterai-research

rohan anil @_arohan_

reply from @_ueaj (ueaj)

rohan anil @_arohan_ · 7h I need to start getting healthier for the singularity. Started lifting very heavy weights. 💬 18 🔁 5 ♥ 130 📊 13K ueaj @_ueaj · 7h Backup plan in case alignment goes wrong and we have to fight Cthulhu in hand to hand combat, obviously
Note from Claude Sonnet 5

Plain text tweet and reply, no images.

singularityhumortwitterai safety

rohan anil @_arohan_

I had one of the best coding sessions last night with claude code. I was trying to do something bit complicated, I started out writing md files with full descriptions of all the gotcha, and how to run, how to debug if you got stuck – this took some effort then I watched Claude nail what otherwise would have taken both expertise & time in just a few hrs. 9:43 AM · Sep 9, 2025 · 1,807 Views
Note from Claude Sonnet 5

A positive testimonial for Claude Code, describing how thorough upfront documentation (md files with gotchas and debug steps) led to strong autonomous performance on a complex coding task. Practical/anecdotal evidence about Claude Code workflow effectiveness.

claude codecoding agentstwitterworkflowai coding assistants

rohan anil @_arohan_

quoting @liang_we... (Weixin Li)

rohan anil @_arohan_ · 4h This is really cool work! I wonder if we could generalize even better by introducing modality as feature embedding to the router instead. That is router gets privileged information. > QUOTED: Weixin Li... @liang_we... · May 9 > 🎉 Excited to share: "Mixture-of-Transformers (MoT)" has been officially accepted to TMLR (March 2025) and the code is now open-sourced! > ... > Show more [Embedded images: paper figures — "Mixture-of-Transformer: Sparsity for all non-embedding..." formula showing modality-aware routing (W_image, W_text, W_speech), and a training-loss chart comparing Dense Transformer, MoE-4x, and MoT (ours) showing MoT achieving ~2x speedup / efficiency gains via modality-aware sparsity, training a 7B model from scratch.]
Note from Claude Sonnet 5

A technical ML Twitter exchange about the "Mixture-of-Transformers" (MoT) paper (accepted TMLR March 2025), a modality-aware sparse architecture showing ~2x training efficiency gains over dense transformers, with a follow-up suggestion to route based on modality embeddings. General ML architecture research, not directly safety-related but relevant to Nathan's technical tracking of model architecture trends (adjacent to his own brain_graph_1 work on routing/sparsity).

twittermachine-learningmixture-of-transformersmodel-architecturesparsitymoerohan-anil