← All topics

generalization

3 captures, most recent first.

François Chollet @fchollet

— saved image

François Chollet @fchollet
There are essentially two main options to remedy this:

1. Find ways to perform active inference, so that the model adapts its learned program in contact with a new data distribution at test time. Would likely lead to some meaningful progress, but it isn't the ultimate solution, more of an incremental improvement.

2. Change the training mechanism to something more robust than SGD, such as the MDL principle. This would pretty much require moving away from deep learning (curve fitting) altogether and embracing discrete program search instead (which I have advocated for many years as a way to tackle reasoning problems...)

2:14 AM · Mar 8, 2024 · 150.5K Views
Note from Claude Sonnet 5

Tweet from François Chollet (@fchollet), dated March 8 2024, proposing two remedies for a generalization problem in deep learning he'd described earlier in the thread: test-time active inference (incremental) or replacing SGD with something like the MDL principle via discrete program search (a bigger departure from curve-fitting deep learning).

deep learning theorygeneralizationprogram searchfrancois chollet

kalomaze @kalomaze

— saved image

kalomaze @kalomaze · 50m
periodic functions seem like the most fundamental class that conventional nonlinearities can't generalize over, even with depth
even tho you have exponentially many piecewise linear regions with ReLU-likes, that's not sufficient to extrapolate a sine wave...
Note from Claude Sonnet 5

A tweet from kalomaze arguing that periodic functions (like sine waves) are a fundamental class conventional neural network nonlinearities (ReLU-likes) can't extrapolate/generalize over, even with depth or many piecewise-linear regions.

neural networksnonlinearitiesgeneralizationperiodic functionsml theorytwitter

rohan anil @_arohan_

rohan anil ✓ 🟦 @_arohan_ · 1h Intelligence requires choosing the right generalization when the datasets underdetermines the world. One of those is the right sets of inductive bias in the learning algorithm, not from compression alone.
Note from Claude Sonnet 5

Text-only post, cut off at bottom edge (engagement icons partially visible but counts not legible); no images.

machine-learninggeneralizationinductive-biastwitterai-research