← All topics

diffusion-models

2 captures, most recent first.

Artur Chakhvadze @norpadon

quoting @jxmnop (jxmo)

Artur Chakhvadze @norpadon · 22h A trick I came up with in ~2017 in the context of image matting is that you can train a network to predict d_loss/d_y instead of y and do a gradient descent for multiple iterations during inference. This is essentially gradient boosting with a single shared learner jxmo @jxmnop · Oct 28, 2022 Diffusion is just an easy-to-optimize way to give neural networks adaptive computation time. Makes sense then that diffusion mode... [cut off]
Note from Claude Sonnet 5

A tweet on an architectural trick — training a network to predict the gradient of the loss (d_loss/d_y) and applying iterative gradient descent at inference, framed as "gradient boosting with a single shared learner" — quoting an older tweet framing diffusion models as adaptive computation time. Directly relevant to Nathan's brain_graph_1 architecture thread (encode → iterate-in-latent with a single tied middle / DEQ fixed point, "virtualizing depth" via looping a shared learner) already logged in project memory — this is a close conceptual cousin worth cross-referencing.

twittermachine-learningarchitecturediffusion-modelsadaptive-computationgradient-boostingdeqinterpretability

Simo Ryu @cloneofsimo

I really want to call out examples, but last time I did that I literally got threats based on 'their research networks and future hire rejections' so I decided to not call out names on twitter but to give you couterexamples: Karras in his paper comes up with (I swear to god) > 4 novel methods and just casually name the paper "Analyzing and Improving" Another GOATs, "Simple Diffusion", "Scalable Shampoo", "Improving Diffusion Models", "Diffusion Beats Gans" They just contribute without cringingly renaming and trying to steal previous work's name [Four screenshots of paper title/abstract pages: (1) Karras et al., NVIDIA/Aalto, "Analyzing and Improving..." diffusion paper, 20 Mar 2024; (2) a diffusion paper abstract with example generated images of a frog, an owl playing piano, and robots; (3) Anil, Gupta, Koren, Singer et al. (Google Research/Tel Aviv/Princeton), "Scalable Shampoo" second-order optimization paper, March 8 2021; (4) Nichol & Dhariwal, "Improved Denoising Diffusion Probabilistic Models", 18 Feb 2021]
Note from Claude Sonnet 5

A machine-learning Twitter post complaining about researchers who inflate paper titles/rename existing methods for credit, contrasted with well-known "GOAT" papers that undersell themselves with plain titles. General ML culture commentary, not directly about AI safety.

twittermachine-learningdiffusion-modelsresearch-cultureacademic-norms