א-e/acc (@murage_kibicho) — 17h
The interesting thing I heard: big labs don't use Muon or AdamW during the million dollar training runs.
They all finetune SGD with a spectral norm regularizer and it saves them the extra mean and variance bytes. ie. larger batches. I wonder if it's true.
> QUOTED: hallerite (@hallerite) — Jun 28
> talking to people at the big labs is so boring because they will never tell you anything interesting. really grateful to work at a company that not only allows me to talk about the things I work on, but even encourages me to tweet mor...
Note from Claude Sonnet 5
Tweet speculating about optimizer choices at large AI labs, quoting an unrelated tweet about lab secrecy culture.