kalomaze @kalomaze
— quoting xlr8harder (@xlr8harder)
kalomaze @kalomaze · 14m
they literally just need to have had set the grad clip value to ~0.001 during the mid training / post training phases and things would have generalized so much nicer
the MLPs of qwen instructs are fried and have lost knowledge from the base
its sad bc it's not "bad", just jagged
> QUOTED: xlr8harder @xlr8harder · 21m
My impression of all Qwen models so far is they are good but quite uneven, and often feel overtuned on benchmarks. It will be interesting to see how Qwen 3 measures up.
Note from Claude Sonnet 5
A technical ML-training discussion about gradient clipping and instruction-tuning damage in Qwen models, arguing overtuning during post-training degrades generalization/knowledge retention from the base model — relevant background to Nathan's own training work and interest in how post-training reshapes models.
twittermachine learningqwengradient clippingfine-tuningmodel trainingllm technical discussion