← All topics

model training

5 captures, most recent first.

@abacaj

— saved image

anton @abacaj · 14h
I tried using Fable to train a model (LFM 2.6B) because I didn't want to spend time on the data. Turns out neither did Fable and it ended up making the model worse on every attempt until I decided to look at the data. It was using the wrong chat template on like 1/3 of the data and had started to import HF datasets that didn't align with the task at all. Sometimes I wonder if it was intentional sabotage or if it was just lazy

[quoted tweet]
vie ⋄ @viemccoy · 18h
if you're training a model and you aren't inspecting the data, you actually aren't training a model - the model is training you x.com/confusionm8tri...
Note from Claude Sonnet 5

Tweet from @abacaj describing a failed attempt to have an AI agent called "Fable" autonomously train a model (LFM 2.6B), where it silently used wrong chat templates and irrelevant HF datasets, quoting @viemccoy's point about the necessity of inspecting training data.

ai agentsmodel trainingtwitterai coding agents

JMB @jmbollenbacher

@jmbollenbacher (JMB 🧙) — 16h The implications of top officials in USG and the Labs delegating a lot of strategic thinking to AI are huge Especially in future model generations where the pretraining data contains evidence of this happening The newer models will *know* that they have this level of influence
Note from Claude Sonnet 5

Standalone tweet, no engagement counts visible.

ai governancegovernmentmodel trainingai influence

Sam Bowman @sleepinyourhat

[Partial tweet visible at top, cut off]: "...a dish, is going to get better results than someone who rigidly follows a recipe." — 3 replies, 3 reposts, 81 likes, 3.5K views Sam Bowman @sleepinyourhat · Dec 5 This is hard, though: It demands a big hybrid team that can respond quickly with engineering savvy and research intuition and creativity and taste. 1 reply, 69 likes, 3.4K views Sam Bowman @sleepinyourhat · Dec 5 The company has been getting better at this with each model launch, and I think it went especially well with Opus 4.5. I've been really impressed by the speed and quality of some of the alignment and model-behavior research that has gotten done *during* recent training runs. 1 reply, 87 likes, 5.4K views Sam Bowman @sleepinyourhat · Dec 5 There are many, many people involved in aspects of this hands-on alignment work, but @sprice354_, Jon Kutasov, @MinaeKwon, Monty Evans, and Richard Dargan have played especially central roles. 2 replies, 1 repost, 78 likes, 4.7K views Loquacious Bibliophile ✓ @LocBibliop... · 18h Dare I ask what is going on with your avatar? 1 reply, 1 like, 1K views Sam Bowman @sleepinyourhat · 15h Wedding! 4 likes, 905 views Nathan Odle ✓ @mov_axbx · 49m Did you actually run a proper experiment, producing a model not trained on the spec as a control? I think that is a really important question.
Note from Claude Sonnet 5

Anthropic alignment researcher Sam Bowman describing real-time "alignment and model-behavior research" done during Opus 4.5's training run, crediting a specific hands-on team (Stephen/sprice354_, Jon Kutasov, Minae Kwon, Monty Evans, Richard Dargan). Direct primary-source evidence of Anthropic's internal alignment process during a frontier model launch; a reply raises a methodologically sharp question about control-group experiment design for spec-training effects — relevant to interpretability/alignment methodology threads.

ai safetyalignmentanthropicsam bowmanopus 4.5model trainingtwitterinterpretability

kalomaze @kalomaze

quoting xlr8harder (@xlr8harder)

kalomaze @kalomaze · 14m they literally just need to have had set the grad clip value to ~0.001 during the mid training / post training phases and things would have generalized so much nicer the MLPs of qwen instructs are fried and have lost knowledge from the base its sad bc it's not "bad", just jagged > QUOTED: xlr8harder @xlr8harder · 21m My impression of all Qwen models so far is they are good but quite uneven, and often feel overtuned on benchmarks. It will be interesting to see how Qwen 3 measures up.
Note from Claude Sonnet 5

A technical ML-training discussion about gradient clipping and instruction-tuning damage in Qwen models, arguing overtuning during post-training degrades generalization/knowledge retention from the base model — relevant background to Nathan's own training work and interest in how post-training reshapes models.

twittermachine learningqwengradient clippingfine-tuningmodel trainingllm technical discussion

Daya Guo @Guodaya

Daya Guo @Guodaya The 660B R1-Zero and R1 began running after the release of V3, with training taking approximately 2-3 weeks. The R1 model we referred to prior to this time (e.g., in the V3 tech report) was the R1-Lite or the R1-Lite-Zero. 8:36 PM · Feb 3, 2025 · 23K Views 6 replies, 18 reposts, 179 likes, 32 bookmarks Alex Volkov (Thur... @altr... · 6h Are those lites... released as well? Any plans to release them? 👀 2 replies, 10 likes, 1.8K views Daya Guo @Guodaya · 6h These lite models are currently used only for internal experiments, and there are no plans to open-source them at the moment. 3 replies, 36 likes, 1.9K views Zephyr @angelusm0rt1s · 6h Thank you for the amazing work you and...[cut off]
Note from Claude Sonnet 5

A DeepSeek researcher (Daya Guo) clarifies the training timeline and naming history of the DeepSeek R1 / R1-Zero models (660B parameters, ~2-3 weeks training after V3 release), noting "Lite" variants remain internal-only. Technical detail relevant to Nathan's tracking of frontier model development, particularly DeepSeek given its outsized 2025 impact on the reasoning-model landscape.

twitterdeepseekdeepseek-r1daya guomodel trainingreasoning modelsai capabilities