Harsh Bhatt @harshbhatt7585
Harsh Bhatt ✓ @harshbhatt7585 · 23h
I tried moun over AdamW with GRPO, it is far stable than that with moun.
[2 replies, 371 views]
rajan agarwal ✓ @_rajanagarwal · 23h
my intuition here is that muon probably provides strong benefits when the intention is generalized RL, but there's a chance that only works with muon bases just by the argument of loss landscapes
[1 reply, 321 views]
Harsh Bhatt ✓ @harshbhatt7585 · 22h
I tried in high dense grid world with HRM model, In my GRPO experiments, Muon felt noticeably more stable than AdamW. I also think that operator-aware step geometry may help more when the RL objective is broad and general rather than narrowly optimized, But grid-world was very narrowed as cannot be as general as LLMs but still it felt better, let me search the wandb experiment, will drop here if I find.
[1 reply, 2 likes, 95 views]
Harsh Bhatt ✓ @harshbhatt7585 · 22h
*HRM model wa not pretrained LLM
[96 views]
Mingta Kaivo 明塔 ... ✓ @MingtaKa... · 12h
super interesting thread. we've been experimenting with different optimizers for fine-tuning audio models at audiowave and the step size sensitivity is real. curious if anyone has tried muon specifically for RLHF on smaller models (sub-7B) or if the benefits mostly show up at scale?
Note from Claude Sonnet 5
An ML researcher discussion thread comparing the Muon optimizer vs AdamW for GRPO (a policy-gradient RL algorithm used in RLHF-style training), including stability observations across grid-world and audio model fine-tuning experiments. Technical ML-training content.
machine learningoptimizersmuonadamwgrporlhftwitterreinforcement learning