Artur Chakhvadze @norpadon
— quoting @jxmnop (jxmo)
Note from Claude Sonnet 5
A tweet on an architectural trick — training a network to predict the gradient of the loss (d_loss/d_y) and applying iterative gradient descent at inference, framed as "gradient boosting with a single shared learner" — quoting an older tweet framing diffusion models as adaptive computation time. Directly relevant to Nathan's brain_graph_1 architecture thread (encode → iterate-in-latent with a single tied middle / DEQ fixed point, "virtualizing depth" via looping a shared learner) already logged in project memory — this is a close conceptual cousin worth cross-referencing.
twittermachine-learningarchitecturediffusion-modelsadaptive-computationgradient-boostingdeqinterpretability