← All topics

grokking

1 capture, most recent first.

Yuandong Tian @tydsh

quoting Jamie Simon (@learning_mech)

Yuandong Tian @tydsh · 2h History repeats itself 😀 The concept "learning mechanics" is not a new thing but actually has been explored for a very long time. It is human nature to think deeper than a blind belief of the scaling laws. I have been working on rigorously modeling training dynamics of deep nonlinear models for many years, with many non-trivial solvable examples in nonlinear dynamics that may be interesting for @learning_mech to take a look. This includes 1. Contrastive learning (e.g., arxiv.org/abs/2110.09348 arxiv.org/abs/2201.12680, arxiv.org/abs/2206.01342) 2. Non-contrastive learning (e.g., arxiv.org/abs/2102.06810, arxiv.org/abs/2110.04947) 3. Training dynamics in Transformers (e.g., arxiv.org/abs/2310.00535, arxiv.org/abs/2305.16380) 4. Grokking behaviors (arxiv.org/abs/2509.21519) 5. Spontaneous symmetry breaking (arxiv.org/abs/1703.00560) 6. Mechanism in forming symbolic solutions from gradient descent (arxiv.org/abs/2410.01779) I am the first/solo author in most of the works listed above. Code are here: github.com/yuandong-tian/... > QUOTED: > Jamie Simon @learning_mech · Apr 24 > 1/ Deep learning is going to have a scientific theory. We can see the pieces starting to come together, and it's looking a lot like physics! > > We're releasing a paper pulling together these ... > > [Table, "Section | Approach | Examples in deep learning | Examples from physics": > 2.1 solvable settings — deep linear networks, kernel regression, multi-index models — harmonic oscillator, hydrogen atom, Ising model > 2.2 simplifying limits — lazy vs. rich learning, width/depth → ∞, small initialization — thermodynamic limit (n,V→∞), classical limit (ħ→0), hydrodynamic limit (k,ω→0) > (table cut off, more rows below)]
Note from Claude Sonnet 5

A machine-learning theory thread: Yuandong Tian responds to Jamie Simon's paper framing deep learning as converging toward a physics-like scientific theory ("learning mechanics"), listing his own prior work on training dynamics, contrastive/non-contrastive learning, grokking, and symmetry breaking as precedent. General ML-theory reading, not directly tied to project's core AI-safety/model-welfare threads but relevant to Nathan's broader ML research interests (brain_graph_1 training dynamics).

machine learninglearning theorytraining dynamicsscaling lawsgrokkingtwitteryuandong tianjamie simon