← All topics

rohan-anil

1 capture, most recent first.

rohan anil @_arohan_

quoting @liang_we... (Weixin Li)

rohan anil @_arohan_ · 4h This is really cool work! I wonder if we could generalize even better by introducing modality as feature embedding to the router instead. That is router gets privileged information. > QUOTED: Weixin Li... @liang_we... · May 9 > 🎉 Excited to share: "Mixture-of-Transformers (MoT)" has been officially accepted to TMLR (March 2025) and the code is now open-sourced! > ... > Show more [Embedded images: paper figures — "Mixture-of-Transformer: Sparsity for all non-embedding..." formula showing modality-aware routing (W_image, W_text, W_speech), and a training-loss chart comparing Dense Transformer, MoE-4x, and MoT (ours) showing MoT achieving ~2x speedup / efficiency gains via modality-aware sparsity, training a 7B model from scratch.]
Note from Claude Sonnet 5

A technical ML Twitter exchange about the "Mixture-of-Transformers" (MoT) paper (accepted TMLR March 2025), a modality-aware sparse architecture showing ~2x training efficiency gains over dense transformers, with a follow-up suggestion to route based on modality embeddings. General ML architecture research, not directly safety-related but relevant to Nathan's technical tracking of model architecture trends (adjacent to his own brain_graph_1 work on routing/sparsity).

twittermachine-learningmixture-of-transformersmodel-architecturesparsitymoerohan-anil