← All topics

gradient descent

2 captures, most recent first.

Ben Grimmer @prof_grimmer

reposted by Yushun Zhang — saved image

Yushun Zhang reposted
Ben Grimmer @prof_grimmer · 5h
A new paper by Jianhao Ma and Yuxin Chen, with proof "developed by GPT-5.6 Sol Pro", answers a question I have cared about for the past few years.

They showed that no gradient descent stepsize schedule (fractally or otherwise) can get full acceleration, i.e., matching Nesterov.
Note from Claude Sonnet 5

Tweet from a math/optimization professor highlighting a new paper by Jianhao Ma and Yuxin Chen, whose proof was 'developed by GPT-5.6 Sol Pro', showing no gradient descent stepsize schedule can achieve full Nesterov-matching acceleration.

optimization theorygradient descentai-assisted proofgpt-5.6twitter

Taelin @VictorTaelin

Taelin ✓ @VictorTaelin Amazing questions, thanks. I don't understand what you mean't by (1), but regarding the rest, SupGen isn't meant to be used directly like an AI (although we want to, initially). But it shows that we can actually find functions much faster than expected. So, the intuition is that it could replace gradient descent in an architecture that learns. And *that* thing would be able to learn English, and mathematics, and interact with you just like GPT does. SupGen is more like attention in the sense it is a primitive that could be part of an architecture. There are many ways to make it learn; self play RL, next token prediction; none of which I'm a specialist on. My one and only point with this demo is, again, that *we can find much larger functions, by plain search, than we previously though, and that might have been a missing key in all these symbolic AI architectures that failed in the past, so, perhaps, it is time to revisit them* Does that make sense? 3:58 PM · Jan 22, 2025 · 1,269 Views [4 replies, 1 repost, 34 likes, 3 bookmarks]
Note from Claude Sonnet 5

Victor Taelin (HVM/Bend language creator) explaining "SupGen," a program-search primitive that finds larger functions via plain search than expected, and speculating it could replace gradient descent as a learning mechanism — a revival-of-symbolic-AI argument. Technical ML architecture discussion.

machine learningprogram synthesissymbolic aigradient descentvictor taelintwitterai architecture