François Fleuret @francoisfleuret
François Fleuret @francoisfleuret
BTW if you are short on memory and have a huge vocabulary, note that materializing the logits matrix may cost you a lot (seqlen 8k, batch size 4, voc 128k, fp16 -> 8Gb)
Fusing the linear readout and the cross-entropy avoids that.
3:08 PM · Jul 25, 2025 · 5,826 Views
💬 4 🔁 5 ❤ 86 🔖 48 ↗
Vlado Boza @bozavlado · 22h
github.com/apple/ml-cross... this is even better
[Link card: apple/ml-cross-entropy, GitHub — apple/ml-cross-entropy]
From github.com
💬 1 🔁 ❤ 12 📊 371 ↗
Max Vox (fka Duke Z...) @DukeZ... · 21h
i would simply design the triaxilated deflector shield geometry to absorb the [cut off]
[Banner overlay: "New tab opened / SWITCH"]
Note from Claude Sonnet 5
A technical ML-engineering thread on memory optimization for training with large vocabularies — fusing linear readout with cross-entropy loss to avoid materializing the full logits matrix, referencing Apple's ml-cross-entropy library. Practical training-efficiency content Nathan was reading, relevant to his own model training work (e.g. brain_graph_1).
twittermachine-learningtraining-efficiencycross-entropymemory-optimizationapple