← All topics

cross-entropy

1 capture, most recent first.

François Fleuret @francoisfleuret

François Fleuret @francoisfleuret BTW if you are short on memory and have a huge vocabulary, note that materializing the logits matrix may cost you a lot (seqlen 8k, batch size 4, voc 128k, fp16 -> 8Gb) Fusing the linear readout and the cross-entropy avoids that. 3:08 PM · Jul 25, 2025 · 5,826 Views 💬 4 🔁 5 ❤ 86 🔖 48 ↗ Vlado Boza @bozavlado · 22h github.com/apple/ml-cross... this is even better [Link card: apple/ml-cross-entropy, GitHub — apple/ml-cross-entropy] From github.com 💬 1 🔁 ❤ 12 📊 371 ↗ Max Vox (fka Duke Z...) @DukeZ... · 21h i would simply design the triaxilated deflector shield geometry to absorb the [cut off] [Banner overlay: "New tab opened / SWITCH"]
Note from Claude Sonnet 5

A technical ML-engineering thread on memory optimization for training with large vocabularies — fusing linear readout with cross-entropy loss to avoid materializing the full logits matrix, referencing Apple's ml-cross-entropy library. Practical training-efficiency content Nathan was reading, relevant to his own model training work (e.g. brain_graph_1).

twittermachine-learningtraining-efficiencycross-entropymemory-optimizationapple