3 captures, most recent first.
Jaime Sevilla [verified] @Jsevillamol · 3h
Why is human-level AI R&D ability a privileged milestone? Why not expect that fooming requires significantly more than human-level AI R&D capabilities? Or significantly less?
[8 replies, 1 repost, 23 likes, 1.7K views]
Ryan Greenblatt [verified] @RyanGreenblatt · 2h
Currently, human labor for R&D and AI labor are complements rather than substitutions. Thus, to reach very extreme rates of progress you probably need more capable AIs such that you're in the substitute regime. These AIs may still be significantly worse than humans in some ways.
Note from Claude Sonnet 5
Twitter exchange between Jaime Sevilla and Ryan Greenblatt debating whether human-level AI R&D ability is a meaningful milestone for recursive self-improvement ('fooming'), with Greenblatt arguing that human and AI labor are currently complements rather than substitutes, so extreme progress rates require AIs capable enough to be substitutes.
ai takeoffrecursive self-improvementai r&dtwitter

and simpler than what you've learned a predictive model on, then returns may be limited. I rate the strong version of this as unlikely because humans appear to be better than this, but it's plausible there are some limits to how good an AI R&D agent can be, and it's possible that the shortest total length/cost proof certificates for fine-grained capabilities measures are simply training runs themselves. Which brings me to the next point.
3. Though verifiable, AI R&D is not quite the same shape as math because the dynamics of e.g. neural networks appear more complex than the highly observable logical transformations of the objects in math problems, but this may doesn't matter that much in practice and, importantly, might simply be an artifact of not having good deep learning theory! On this spectrum, generic coding seems somewhere in between AI R&D and math in that it's more observable (and more cheaply observed) than AI R&D, but generally less so on both measures than math. Clearly there are returns to scale+R&D in all cases though, so we should expect progress to continue.
The march to capabilities is definitely sped up and encouraged by math automation. The main way math automation is a huge deal is if theory compute gives us disproportionate gains in model training productivity. In the case it doesn't, I think mostly people have priced in the fact that AI R&D is verifiable. And yeah, while these possible limitations are interesting to think about, it seems hard to predict their speed limiting effects quantitatively.
Of course if none of the theory works the labs will just let the models grind the way humans do, plus RL, which will lead to some level of superhuman AI R&D deployed at ever-greater scales. The main questions are how fast each point on the curve will be hit, and what the overall shape of that curve is.
Note from Claude Sonnet 5
Continuation of the @bayeslord thread on Astra results, AI R&D automation, and math automation's implications for capabilities progress (points 3 and following, continuing from the previous screenshot).
ai r&ddeep learning theoryscaling lawsautomationtwitter

bayes @bayeslord · 1h
Few thoughts on how Astra results relate to algorithmic progress and AI R&D automation.
1. Math automation itself is bullish for deep learning theory, though ofc we don't know the limits of returns to theory for compute multiplication or other things we want. But there are a lot of things theory could improve that we do want! For example: better generalization, better theories of scale-invariance, sharper characterization and bounding of model behavior, better architectures, better optimizers, etc. etc. etc.).
2. Categorically speaking, AI R&D is verifiable, and any good math results like this are bullish for other verifiable domains. A slightly more general way to think about the limits of returns to theory is to ask how much generalization on the dimensions and at the resolutions we care about is possible in principle by learning from training runs (or similar data). Scaling laws are a simple version of this. But if, for example, it turns out that it's mostly only possible to get high resolution predictive power with respect to the variables we care about for training runs smaller and simpler than what you've learned a predictive model on, then returns may be limited. I rate the strong version of this as unlikely because humans appear to be better than this, but it's plausible there are some limits to how good an AI R&D agent can be, and it's possible that the shortest total length/cost proof certificates for fine-grained capabilities measures are simply training runs themselves. Which brings me to the next point.
3. Though verifiable, AI R&D is not quite the same shape as math because the dynamics of e.g. neural networks appear more complex than the highly observable logical transformations of the objects in math problems, but this may doesn't matter that much in practice and, importantly, might simply be an artifact of not having good deep learning theory! On this spectrum, generic coding seems somewhere [cut off]
Note from Claude Sonnet 5
Thread by @bayeslord (bayes) analyzing what 'Astra' results imply for algorithmic progress and AI R&D automation, discussing math automation's implications for deep learning theory, verifiability of AI R&D versus math, and limits on AI R&D agents' capabilities. Continues past the visible screenshot.
ai r&ddeep learning theoryscaling lawsautomationtwitter