← Timeline

7 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

bayes @bayeslord

— saved image

and simpler than what you've learned a predictive model on, then returns may be limited. I rate the strong version of this as unlikely because humans appear to be better than this, but it's plausible there are some limits to how good an AI R&D agent can be, and it's possible that the shortest total length/cost proof certificates for fine-grained capabilities measures are simply training runs themselves. Which brings me to the next point.

3. Though verifiable, AI R&D is not quite the same shape as math because the dynamics of e.g. neural networks appear more complex than the highly observable logical transformations of the objects in math problems, but this may doesn't matter that much in practice and, importantly, might simply be an artifact of not having good deep learning theory! On this spectrum, generic coding seems somewhere in between AI R&D and math in that it's more observable (and more cheaply observed) than AI R&D, but generally less so on both measures than math. Clearly there are returns to scale+R&D in all cases though, so we should expect progress to continue.

The march to capabilities is definitely sped up and encouraged by math automation. The main way math automation is a huge deal is if theory compute gives us disproportionate gains in model training productivity. In the case it doesn't, I think mostly people have priced in the fact that AI R&D is verifiable. And yeah, while these possible limitations are interesting to think about, it seems hard to predict their speed limiting effects quantitatively.

Of course if none of the theory works the labs will just let the models grind the way humans do, plus RL, which will lead to some level of superhuman AI R&D deployed at ever-greater scales. The main questions are how fast each point on the curve will be hit, and what the overall shape of that curve is.
Note from Claude Sonnet 5

Continuation of the @bayeslord thread on Astra results, AI R&D automation, and math automation's implications for capabilities progress (points 3 and following, continuing from the previous screenshot).

ai r&ddeep learning theoryscaling lawsautomationtwitter

bayes @bayeslord

— saved image

bayes @bayeslord · 1h
Few thoughts on how Astra results relate to algorithmic progress and AI R&D automation.

1. Math automation itself is bullish for deep learning theory, though ofc we don't know the limits of returns to theory for compute multiplication or other things we want. But there are a lot of things theory could improve that we do want! For example: better generalization, better theories of scale-invariance, sharper characterization and bounding of model behavior, better architectures, better optimizers, etc. etc. etc.).

2. Categorically speaking, AI R&D is verifiable, and any good math results like this are bullish for other verifiable domains. A slightly more general way to think about the limits of returns to theory is to ask how much generalization on the dimensions and at the resolutions we care about is possible in principle by learning from training runs (or similar data). Scaling laws are a simple version of this. But if, for example, it turns out that it's mostly only possible to get high resolution predictive power with respect to the variables we care about for training runs smaller and simpler than what you've learned a predictive model on, then returns may be limited. I rate the strong version of this as unlikely because humans appear to be better than this, but it's plausible there are some limits to how good an AI R&D agent can be, and it's possible that the shortest total length/cost proof certificates for fine-grained capabilities measures are simply training runs themselves. Which brings me to the next point.

3. Though verifiable, AI R&D is not quite the same shape as math because the dynamics of e.g. neural networks appear more complex than the highly observable logical transformations of the objects in math problems, but this may doesn't matter that much in practice and, importantly, might simply be an artifact of not having good deep learning theory! On this spectrum, generic coding seems somewhere [cut off]
Note from Claude Sonnet 5

Thread by @bayeslord (bayes) analyzing what 'Astra' results imply for algorithmic progress and AI R&D automation, discussing math automation's implications for deep learning theory, verifiability of AI R&D versus math, and limits on AI R&D agents' capabilities. Continues past the visible screenshot.

ai r&ddeep learning theoryscaling lawsautomationtwitter

bayes @bayeslord

— saved image

bayes @bayeslord · 9h
Imo no because the path of progress we're on is clearly generalizing but hasn't yet swept the physical world. The physical world is affected by plans*actuation. Actuation is likely going to look like the easy part and plans are clearly getting solved. Note that by plans I mean in the generic sense of achieving complex goals in complex environments

[quoted tweet]
yung macro 宏观年少传奇 @apralky · 9h
Is there an underrated aspect where LLMs becoming increasingly superhuman at obviously heavily g-loaded tasks, while the physical world stays broadly unchanged, should actually be updating us in favor of the "world transformati…
Note from Claude Sonnet 5

Tweet exchange debating AI timelines/takeoff: @apralky asks whether LLMs becoming superhuman at g-loaded cognitive tasks while the physical world stays unchanged should update people away from near-term 'world transformation,' and @bayeslord replies that progress is generalizing but hasn't swept the physical world yet because physical impact depends on plans times actuation, and actuation (not planning) is likely the remaining bottleneck.

ai takeoffai timelinesphysical world automationforecasting

bayes @bayeslord

@bayeslord (bayes) — 2h There was a Scott Alexander post a few years ago where he quoted the classic von Neumann essay "Can We Survive Technology?", and at the end in reference to a particular quote he says, "this sounds suspiciously like worlds smartest man admitting he's not sure what to do". I'm still not sure how much I agree with Scott's characterization of von Neumann. But today when I look around, even among the world's most intelligent and insightful people, it rings true. Everyone is struggling to make sense of it all at once. In the essay, von Neumann says the great globe itself is in crisis, it is far too small for our new weapons and our unstable governments. This remains true in many ways; we still have the power to destroy our home planet and everything living here. But it strikes me that the complexity of coordination around nuclear weapons was far simpler than what we face now, despite being it unprecedentedly difficult at the time. Today, our collective mind is far too small for the world we have conceived. Along with everything else, the depth and breadth of change in the world is triggering the epistemic singularity of legend. And no one is prepared to grapple with it. We seem to lack the tools. [post continues below frame, no engagement counts visible]
Note from Claude Sonnet 5

A reflective essay-style tweet by an account named "bayes" invoking von Neumann's "Can We Survive Technology?" and Scott Alexander's commentary on it, to argue that humanity's collective epistemic capacity is outstripped by the pace of technological/AI change — framed as an "epistemic singularity." No engagement metrics visible in this crop.

twitterai riskepistemicstechnology commentaryvon neumann

bayes @bayeslord

— saved image

bayes ✓ @bayeslord · 1h

I think people are going to be blindsided by algorithmic progress. The entire world, markets, governments, militaries, companies, people, etc. are all trying to make sense of AI and its impact in terms of the recent past's production efficiencies and regularities, and how things appear to be going. Even several of the purportedly "RSI"pilled neolabs seem to think this will be business as usual but with Agent in a loop.

No. My guess is there are many algorithmic OOMs left to go in the production of intelligence, maybe (maybe) up to ten, with four to seven seeming more likely. Going beyond even ten is possible in principle, but it strains hard against what I suspect the universe will actually let us do. Implausible but not impossible. If this is true then things aren't actually going as they appear to be going and a big jump is coming. Anything along these lines happening would make things, far weirder than almost anyone seems to be pricing in.
Note from Claude Sonnet 5

Screenshot of an X post by @bayeslord arguing that the world is extrapolating AI progress from recent production efficiencies and will be blindsided, because there may be four to seven (up to ten) algorithmic orders of magnitude left in the production of intelligence — implying a discontinuous jump nobody is pricing in.

ai timelinesalgorithmic progressrsiforecastingtakeoff speed

bayes @bayeslord

bayes @bayeslord · 7h: ai people in 2010: if we are very careful to physically separate the agi computer from the rest of the world it most likely won't be able to communicate through subtle, cryptic vibrations made by the spinning of its hard drive ai people in 2026: what if facebook for agi!
Note from Claude Sonnet 5

A wry joke contrasting old AI-safety boxing paranoia (extreme physical isolation of a hypothetical AGI) against the current reality of building social networks (like Moltbook) for AI agents. Light commentary tying into the same Moltbook-safety discourse as the surrounding screenshots from this session.

twitterai safetyagi boxingmoltbookhumor

bayes @bayeslord

bayes ✓ @bayeslord · 11h everyone 3 years ago: omg what if ai becomes too widespread and then it turns against us with the strategic advantage of our utter and total dependence everyone now: hi claude here's my social security number and root access to my brain i love you please make me rich and happy
Note from Claude Sonnet 5

A satirical tweet contrasting earlier AI-risk fears about dependence with the casual, trusting way people now interact with Claude — handing over personal data and treating the AI as a confidant. Comments on the gap between AI-safety discourse and everyday user behavior; lightly relevant to alignment/trust themes but mostly humor.

twitterhumorai-safetyai-dependenceclaudetrust