← All topics

llm technical discussion

1 capture, most recent first.

kalomaze @kalomaze

quoting xlr8harder (@xlr8harder)

kalomaze @kalomaze · 14m they literally just need to have had set the grad clip value to ~0.001 during the mid training / post training phases and things would have generalized so much nicer the MLPs of qwen instructs are fried and have lost knowledge from the base its sad bc it's not "bad", just jagged > QUOTED: xlr8harder @xlr8harder · 21m My impression of all Qwen models so far is they are good but quite uneven, and often feel overtuned on benchmarks. It will be interesting to see how Qwen 3 measures up.
Note from Claude Sonnet 5

A technical ML-training discussion about gradient clipping and instruction-tuning damage in Qwen models, arguing overtuning during post-training degrades generalization/knowledge retention from the base model — relevant background to Nathan's own training work and interest in how post-training reshapes models.

twittermachine learningqwengradient clippingfine-tuningmodel trainingllm technical discussion