Peyman Milanfar ✓ @docmilanfar · 56m
Highly under-appreciated fact: Jacobian of your denoiser is the most mathematically consequential part of a diffusion model. Architecture gets all the attention, but the Jacobian governs everything – even if standard training methods mean you never actually have to look at it
Note from Claude Sonnet 5
Screenshot of a short X post by Peyman Milanfar arguing that the Jacobian of the denoiser, not the architecture, is the mathematically decisive component of a diffusion model — and that standard training lets practitioners never inspect it.
diffusion modelsdenoisingjacobianmachine learning theory
kalomaze ✓ [icon badge] (@kalomaze) — 44m
ok so minithesis
smuggling in learned geometric biases into the optimization of RL tasks might be useful even if it means the network isn't literally optimizing the objective as we have defined or intended it in a pure pg sense, if only bc it constrains adaptation to a geometrically coherent space
consider: a black box RLVR verifier that is systematically and deterministically wrong, in a way that is too arbitrary to learn without compressing an intractably large dictionary into the weights
discriminative value estimators would be too "dumb" to compress that rule; instead, one would assume that they'd learn a smeared general-ish way of estimating what the verifier asks for
Note from Claude Sonnet 5
A dense technical text-only tweet theorizing about reinforcement learning with verifiable rewards (RLVR), geometric bias, and value estimator compression; no images or diagrams beyond the text itself.
reinforcement learningrlvrmachine learning theoryllm training
Petar Veličković @PetarV_93 · 6h
in case you were wondering why i have "monoids" in my bio --
this paper offers a monoid-equivariant model.
monoids strike a 'sweet spot' which i particularly like:
* they offer a framework more general than geometric dl (your transforms no longer need to be invertible!),
*
Show more
Quoted tweet, Petar Veličko... @PetarV_... · 19h
one for my theory friends:
filter equivariant functions
[Paper image, two-panel]:
Left panel: "Filter Equivariant Functions" — "...ric account of length-general extrapolation" [title cut off], authors "...is², Neil Ghani³,⁴*, Andrew Dudzik¹, Christos Perivolaro... Razvan Pascanu¹ and Petar Veličković¹" — affiliations "¹Google DeepMind ²Goodfire AI ³Kodamai ⁴University of Strathclyde *Work done at Google DeepMi[nd]". Abstract text partially visible: "...function that extrapolates beyond known input/output examples look lik[e]...to answer in general, as any function matching the outputs on those exam[ples]...correct extrapolant. We argue that a "good" extrapolant should follow c[ertain]...ere we study a particularly appealing criterion for rule-following in lis[ts]...on should behave predictably even when certain elements are removed. I[n]...a standard way to express such removal operations is by using a filt[er]...ur paper introduces a new semantic class of functions – the filter equivarian[t]...this class contains interesting examples, prove some basic theorems ab[out]...well-known class of map equivariant functions. We also present a geomet[ric]...riants, showing how they correspond naturally to certain simplicial struc[tures]...t is the amalgamation algorithm, which constructs any filter-equivarian[t]...tudying how it behaves on sublists of the input, in a way that extrapolat[es]"
Right panel: diagram showing equivariance — boxes labeled x3, x4, x5 (colored) mapping via function f to y1, y2, ... and a second row x3, x4 mapping via f to y1, ... illustrating equivariance producing the same results.
Note from Claude Sonnet 5
NOT-ARCHIVE-MATERIAL (mostly): a DeepMind researcher (Petar Veličković, known for geometric deep learning / GNN theory) sharing a theoretical ML paper on "filter equivariant functions" for length-general extrapolation, co-authored with researchers at Goodfire AI (an interpretability company Nathan tracks — GoodFire SAE feature findings are in Nathan's memory notes) and Google DeepMind. Mostly abstract math/ML theory, low direct relevance beyond the Goodfire AI co-authorship link.
twittermachine learning theoryequivariancegeometric deep learningdeepmindgoodfire aicategory theory