← All topics

ai self-improvement

1 capture, most recent first.

Yo Shavit @yonashav

— saved image

Yo Shavit @yonashav · 9h
I wonder whether we will soon start to see faster AI self-improvement at OpenAI vs. Anthropic based on the former's known deeper investment in RL, TTC, and math proving more useful for tasks related to AI R&D, and that this gap may grow significantly over the next 6 months.

(Obviously possible Ant is seeing similar results w/ internal models, but my weakly-held sense is that they're not.)

If so, it seems *really* crucial for OAI to be able to correctly calibrate its relative position in the RSI ramp so it can incorporate it into its alignment+security decision-making, especially related to full RSI. The Allies rushed to a bomb on the incorrect assumption that the Axis was right on their tail, when in fact they were far behind and the bomb was plausibly unnecessary, purely due to fog of war. We also saw something similar with the original strawberry results, where OpenAI felt an intense sense of urgency based on rumors that Anthropic+GDM were discovering it in parallel a couple months behind when in fact I've now heard they were ~9 months behind on it.

I expect OpenAI would act very differently wrt future concerning misalignment findings if they knew they were 3 months ahead. It would be very simple for the parties to resolve such an uncertainty. Executing on a pace-info-sharing scheme probably not crucial today, but will be in a few months.
Note from Claude Sonnet 5

Tweet from Yo Shavit (OpenAI) speculating that OpenAI may be pulling ahead of Anthropic in AI self-improvement due to greater RL/test-time-compute investment, drawing an analogy to the Manhattan Project's mistaken urgency about Axis nuclear progress, and arguing labs should calibrate their relative competitive position (e.g. via pace-info-sharing) to make better alignment/security decisions around recursive self-improvement (RSI).

openaianthropicai self-improvementrsiai race dynamicsalignmenttwitter