← Timeline

@murage_kibicho

@murage_kibicho on X

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@murage_kibicho

א-e/acc (@murage_kibicho) — 17h The interesting thing I heard: big labs don't use Muon or AdamW during the million dollar training runs. They all finetune SGD with a spectral norm regularizer and it saves them the extra mean and variance bytes. ie. larger batches. I wonder if it's true. > QUOTED: hallerite (@hallerite) — Jun 28 > talking to people at the big labs is so boring because they will never tell you anything interesting. really grateful to work at a company that not only allows me to talk about the things I work on, but even encourages me to tweet mor...
Note from Claude Sonnet 5

Tweet speculating about optimizer choices at large AI labs, quoting an unrelated tweet about lab secrecy culture.

machine learningoptimizersai labstwitter