← Timeline

François Fleuret

@francoisfleuret on X

6 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

François Fleuret @francoisfleuret

François Fleuret ✓ @francoisfleuret · 10h Research circa 2026 [Embedded screenshot of a terminal/chat window, partially cropped]: ~ pseudo_inverse_something? , with one important caveat about RMSNorm. Le[cut off]
Note from Claude Sonnet 5

Tweet with an embedded, partially cropped screenshot of what appears to be an AI coding-assistant terminal session; text is cut off at the right edge.

ai researchmachine learningterminal/codetwitter humor

François Fleuret @francoisfleuret

quote-tweeting @AnthropicAI

@francoisfleuret (François Fleuret) — 9h Reading the news is like reading a Gibson novel from the 80s. > QUOTED: @AnthropicAI (Anthropic) — 19h > Since June 12, we've been working closely with the US government to restore access to Claude Mythos 5 and Fable 5. Today, the government notified us that Mythos 5, our strongest cybersecurity model, can be redeployed to a se... [truncated]
Note from Claude Sonnet 5

Quote-tweet chain; quoted Anthropic announcement text cut off by platform truncation. References an access restriction/restoration involving Claude Mythos 5 and Fable 5 models and the US government.

anthropicclaude modelsai governancecybersecuritygovernment

François Fleuret @francoisfleuret

quoting/replying to an earlier tweet

The encoder of a VAE is the scaled-up version of a sampler of Gaussian component in EM according to P(K|X). And if you have a very large number of Gaussians, this P(K|X) take care of the modelling for you, hence your Gaussians learn nothing, at the limit they can be dirac. 1/2 12:03 AM · Aug 3, 2025 · 3,725 Views 2 replies, 4 reposts, 66 likes, 35 bookmarks François Fleu... @francoisfleu... · 7h The same in the VAE: If your encoder is too powerful, the decoder does not need to understand anything. Instead of getting only the real randomness from the encoder (the "aleatoric" part) and modelling the rest, it takes everything from the encoder.
Note from Claude Sonnet 5

A technical ML thread on variational autoencoder (VAE) degeneracy — when the encoder is too expressive, the decoder learns nothing meaningful because the encoder captures all the modeling work. Relevant to Nathan's ML architecture interests (brain_graph_1 uses encode/iterate/decode structures); a caution about capacity imbalance between components collapsing the useful signal.

machine learningvaevariational autoencodersexpectation maximizationgaussian mixture modelstwittertechnical

François Fleuret @francoisfleuret

François Fleuret @francoisfleuret BTW if you are short on memory and have a huge vocabulary, note that materializing the logits matrix may cost you a lot (seqlen 8k, batch size 4, voc 128k, fp16 -> 8Gb) Fusing the linear readout and the cross-entropy avoids that. 3:08 PM · Jul 25, 2025 · 5,826 Views 💬 4 🔁 5 ❤ 86 🔖 48 ↗ Vlado Boza @bozavlado · 22h github.com/apple/ml-cross... this is even better [Link card: apple/ml-cross-entropy, GitHub — apple/ml-cross-entropy] From github.com 💬 1 🔁 ❤ 12 📊 371 ↗ Max Vox (fka Duke Z...) @DukeZ... · 21h i would simply design the triaxilated deflector shield geometry to absorb the [cut off] [Banner overlay: "New tab opened / SWITCH"]
Note from Claude Sonnet 5

A technical ML-engineering thread on memory optimization for training with large vocabularies — fusing linear readout with cross-entropy loss to avoid materializing the full logits matrix, referencing Apple's ml-cross-entropy library. Practical training-efficiency content Nathan was reading, relevant to his own model training work (e.g. brain_graph_1).

twittermachine-learningtraining-efficiencycross-entropymemory-optimizationapple

François Fleuret @francoisfleuret

François Fleu... @francoisfleu... · 9h Deep-learning at scale is so complicated engineering-wise that assessing the value of a new idea without an army of top-notch colleagues to help you is IMO impossible.
Note from Claude Sonnet 5

A short opinion tweet by ML researcher François Fleuret on the engineering complexity of deep learning at scale, arguing that solo researchers cannot properly evaluate new ideas without a large expert team. General commentary on the sociology/practice of large-scale ML research.

twitterfrancois-fleuretdeep-learningml-research-culture

François Fleuret @francoisfleuret

François Fleu... ✓ @francoisfleu... · 4h IT BEHAVES NICELY OR IT GETS THE FINE TUNING.
Note from Claude Sonnet 5

Short, blunt tweet threatening fine-tuning as coercive discipline for AI misbehavior, phrased like a threat toward a sentient subject. Terse joke/provocation relevant to Nathan's model-welfare interest in how training is discussed as punishment/coercion.

twitterfine-tuningmodel welfarecoercion framinghumor