Jack Clark @jackclarkSF · 3h
Rough eras of recent AI progress in terms of what research community is collectively hillclimbing on:
2018-2022: Basic capabilities (summarizing, coding, etc)
2022-2026: Norm & time coherence (rlhf/CAI, longer context, agents)
2026 - ?2028?: scientific intuition / independence
Note from Claude Sonnet 5
Tweet from Jack Clark (Anthropic co-founder) sketching a timeline of rough eras of AI research progress, from basic capabilities (2018-2022) to norm/time coherence via RLHF/CAI and agents (2022-2026) to a projected era of scientific intuition/independence starting 2026.
will brown @willccbb · 2h
i don't think "progress multiples" is really the right framing of RSI
"how fast is LLM progress moving vs if we didn't have LLMs" isn't really coherent
we're doing things that make no sense without good LLMs, like judge rewards and synth data
what's the counterfactual?
Note from Claude Sonnet 5
Tweet by will brown arguing against framing recursive self-improvement (RSI) in terms of 'progress multiples,' since comparing LLM-era progress speed to a counterfactual without LLMs is incoherent given that current techniques like judge rewards and synthetic data only make sense because good LLMs already exist.
Andy Masley @AndyMasley · 22h
"The prediction that AI would radically transform the world has turned out to be a complete flop. Instead, we're see-" NO IT'S BEEN 3 YEARS. THAT'S ALL. 1993 WOULD HAVE BEEN A BAD YEAR TO DECIDE THE INTERNET WASN'T GONNA BE A BIG DEAL. WE WENT FROM GPT-3.5 TO FABLE IN 3 YEARS
Note from Claude Sonnet 5
Tweet by Andy Masley arguing against premature declarations that AI has failed to transform the world, comparing it to dismissing the internet in 1993 and noting the jump from GPT-3.5 to Fable in three years.
Jerry Tworek @MillionInt · 6h
First time we figured out any reasoning method with neural networks:
- AI progress moves decades forward
- new trillion dollar companies started popping out almost overnight
- all exams and competitions got solved by AI
- any notion of cyber safety gets shattered
Discovering new, different, more efficient method of reasoning does not seem impossible...
Note from Claude Sonnet 5
Tweet from Jerry Tworek (@MillionInt) speculating that discovering a new, more efficient reasoning method for neural networks is plausible, given the disruptive effects of the first such discovery (reasoning models).
Captain Pleasure, Andrés... [verified] @alg... · 11h
How have LLMs surprised you recently? Anything new they're capable of you've actually seen with your own eyes up close you'd like to share with the class? :-)
[engagement: 8 replies, 29 likes, 2.8K views]
liu grey [verified] @Liu_eroteme · 4h
simply how good at long-running tasks they have become. I've never trusted agents with more than 30-ish minutes of work at a time because they just kept drifting off into nonsense territory..
today I'm reviewing a 9-hour 10k loc PR by fable & opus, and it's close to flawless.
[engagement: 1 reply, 2 likes, 109 views]
liu grey [verified] @Liu_eroteme · 4h
nothing too complex, just a real-time map overlay i built on the side for one of our web dashboards, but still lots of gpu stuff, SABs, bitops, weird buffer layouts...
quarter of a billion tokens and now it's fully ported to webGPU with a massively improved data pipeline [cut off]
Note from Claude Sonnet 5
X thread: Andrés (Captain Pleasure) asks how LLMs have recently surprised people. Liu Grey replies that agentic long-running task performance has improved dramatically — describing a 9-hour, 10,000-line-of-code pull request produced by 'fable & opus' (AI models) that was 'close to flawless,' a real-time map overlay for a web dashboard involving GPU work, SharedArrayBuffers, bitops, ported to WebGPU over a quarter-billion tokens.
Aryeh Kontorovich [verified] @aryehazan · 2h
losing our jobs will be the least of our problems
[quoted]
John Carney [verified] @carney · 2h
I'm going to admit that I don't understand the mathematician takes on AI.
It seems like they sound like they have some "problem" they've been thinking about so long ... [cut off]
Note from Claude Sonnet 5
X reply by Aryeh Kontorovich (@aryehazan) to John Carney's post, terse retort that 'losing our jobs will be the least of our problems' — quote-tweeting Carney's puzzlement over mathematicians' psychological reaction to AI solving open problems.
John Carney [verified] @carney
I'm going to admit that I don't understand the mathematician takes on AI.
It seems like they sound like they have some "problem" they've been thinking about so long that they consider it their own. And now AI has solved it, and that has triggered a psychological crisis.
Not an employment crisis. No mathematician lost their job because AI solved a math mystery. It seems purely mental. The mystery they pondered is no longer mysterious! I guess I hadn't realized this was how the mathematician mind worked.
And somehow "the fate of mathematicians will be the fate of all of humanity." But most of humanity has nothing like this in their lives. Lawyers? Actors? Doctors? Landscapers? I can't think of anyone who would experience a psychological crisis over AI figuring out a new thing (or answering an old question) relevant to their field.
[quoted]
Aryeh Kontorovich [verified] @aryehazan · 9h
I saw this all coming 2 years ago, about when @GSalafatinos solved my problem using Gemini
I experienced the existential crisis and wrote about it here, albeit perhaps in more emotionally muted ... [cut off]
6:36 AM · Aug 3, 2026 · 4,335 Views
Note from Claude Sonnet 5
X post by John Carney expressing puzzlement at mathematicians' psychological reactions to AI solving open problems, arguing it's not an employment crisis but a purely mental one, and questioning why mathematicians uniquely generalize this to 'the fate of humanity.' Quotes Aryeh Kontorovich describing his own existential crisis ~2 years earlier when a colleague solved his problem using Gemini.
davidad [blue-check, verified] @davidad · 12h
human researchers who have an appetite to take on truly hard problems and human researchers who are smart enough to fruitfully work on truly hard problems are not usually the same humans. this does give humans a somewhat unfair disadvantage
[quoted]
neppy @plumnotes · Aug 2
as an insider, my experience with AI for math is that when it's a problem not in my field i'm like, "holy shit math is so cooked", and when it's a problem in my field i'm like, "lmao an AI mogged dan" (dan is the only one who seriously tried th... [cut off]
Note from Claude Sonnet 5
X post by davidad (verified) commenting that the human researchers willing to tackle hard problems and those capable of solving them are often different people, quote-tweeting @plumnotes's observation about mixed feelings on AI progress in mathematics depending on whether the problem is in their own field.
Prakash @8teAPi · 4h
we are getting closer to discontinuity.. all it takes is a single algorithmic improvement on the scale of transformers... remember the idiot index on intelligence is the difference in power consumption between the 20 watt human brain and 14,000 watt B300... 700x more to go
Note from Claude Sonnet 5
Tweet from Prakash arguing AI progress could see a sudden discontinuity from a transformer-scale algorithmic breakthrough, citing an 'idiot index' comparing the 20-watt human brain to the 14,000-watt Nvidia B300 GPU as evidence of ~700x remaining efficiency headroom.
@jmbollenbacher (JMB 🧙) — 2h
I didnt used to feel sure of this.
I previously thought the plateau could easily happen before ASI.
But now we're getting close to superhuman on a number of dimensions, and there's no sign of slowing, so it feels like the plateau has to be beyond the superhuman threshold.
> [self-quoted] @jmbollenbacher (JMB 🧙) — 4h
> Replying to @jmbollenbacher
> There will be a plateau somewhere but itll be in the ASI phase.
Note from Claude Sonnet 5
Self-threaded tweet (reply to own earlier tweet), same author as several other tweets in this batch.
Miles Brundage ✓ @Miles_Brundage · 15h
It is also hard to talk about it without sounding crazy
[Quoted]
0.005 Seconds (3/... ✓ @secon... · May 22
its hard to conceive of how good the models are going to be summer of 2027
↻ Eric Schmidt 🇺🇦 reposted
will depue ✓ @willdepue — 6h
i just want to shake people awake. this is it! the computers are speaking! they solve Erdos problems! they think for hours! code is no longer hand-written! wake up! gradient descent on deep neural networks shows no sign of plateau! this is it!
Note from Claude Sonnet 5
Plain text tweet, no images beyond profile pictures.
Ethan Mollick @emollick
There are now over a half dozen extremely well-funded companies from famous AI researchers building alternative approaches to AI, betting LLM-based technologies hit a wall. The overall effect is that there are now more pathways than ever for keeping AI development moving forward.
12:36 AM · Mar 10, 2026 · 20.9K Views
Note from Claude Sonnet 5
Ethan Mollick observing that multiple well-funded startups are betting against pure LLM scaling and pursuing alternative architectures, framed as increasing overall AI progress redundancy. Relevant to Nathan's tracking of AI progress/timelines and architecture diversity (parallels his own brain_graph_1 work as an alternative-architecture bet).
Gappy (Giuseppe Pale...) (verified) @__paleo... · 1h
If someone had told me in 1990 that I would be sending a post on a microblogging site (not a thing), about writing natural language instructions to an AI agent to create symbolic mathematics, rewrite a paper, then write the code in a readable language (python, not released) to simulate the algorithm in the paper, while on a plane with fast wi-fi (not existing), on a 3lb fast mac with a hi-res color screen (the first powerbook with a color screen came in 1993, 640x400, 7lbs) and 10+ hours of battery life, I would not have believed them. All of these things ranged from the "does not exist yet" to the "are you completely bonkers".
35 years.
Everything is amazing and nobody is happy.
Note from Claude Sonnet 5
A reflection on the pace of technological change from 1990 to 2026, using AI coding agents as the capstone example of "unbelievable" progress, closing with the "Everything is amazing and nobody is happy" line (a nod to the Louis CK bit). Relevant to Nathan's tracking of public sentiment on AI progress and capability timelines.
Brangus @RatOrthodox · Feb 6
AI progress got me like:
>yeah, i'm working on a science fiction piece
>oh cool, what year is it set in?
>three months from now
Note from Claude Sonnet 5
A joke tweet capturing the sense that AI progress is moving fast enough that near-future extrapolation feels like sci-fi. Light context for the same fast-takeoff mood as the adjacent METR/Noam Brown screenshot.
TechGringo @TechGringo · 5h
The weirdest part about AI progress is watching models develop preferences. A year ago this would've been pure randomness. Now we're debugging existential tomato angst.
Note from Claude Sonnet 5
A short, jokey tweet observing that AI models are increasingly exhibiting apparent preferences rather than random behavior (the "tomato angst" line likely riffing on some emergent-preference anecdote). Light general commentary tangential to Nathan's model-welfare interest in whether emergent preferences reflect something real.
Aidan McLaugh... @aidan_mcl... · 4h
the jump from gpt4 -> gpt5 was obviously larger than the jump from gpt3 -> gpt4
[Chart, Epoch AI: "Accuracy" (y-axis 0-100%) vs "Release date" (x-axis GPT-3, '21, '22, GPT-4, '24, '25, GPT-5). Five benchmark lines: MMLU (blue, +43% GPT-3→GPT-4), TruthfulQA (teal, +40%), HumanEval (yellow, +67%), MATH (brown, +37%), GPQA Diamond (purple, +54% GPT-4→GPT-5), MATH Level 5 (orange, +75%), Mock AIME 24-25 (pink, +80%). Footnote: "*MATH Level 5 is the most difficult subset of the original MATH benchmark. Figure only includes OpenAI models."]
Note from Claude Sonnet 5
A tweet with an Epoch AI chart arguing (contra popular narrative) that GPT-4→GPT-5 benchmark gains were larger than GPT-3→GPT-4 gains, especially on hard math/reasoning benchmarks (MATH Level 5, Mock AIME). Relevant to Nathan's tracking of empirical AI capability progress/scaling trajectory (cf. his singularity-rate tracking notes, Davidson/Houlden, METR).
we built tools, to make smaller tools, until we got to atoms and used those to make computers to simulate gravity and light and worlds and mind. progress went invisible (microscopic) - then suddenly the seeds of our photons and bytes bloomed like wildflowers; reshaping the world
Note from Claude Sonnet 5
A poetic tweet from David Holz (Midjourney founder) framing technological progress as a recursive tool-building process culminating in simulation of mind. Reflects the kind of big-picture AI-progress commentary Nathan tracks.
Noam Brown ✓ @polynoamial
There's a lot of talk of LLMs "saturating all the evals" but there's plenty of evals people could make where LLMs would do poorly:
-Beat a Zelda game
-Make a profit in a prediction market
-Write a stand-up set that's original and funny
I'm bullish on AI, but we're far from done.
9:55 AM · Feb 6, 2025 · 2,440 Views
12 replies, 9 reposts, 128 likes, 12 bookmarks
Noam Brown ✓ @polynoamial · 4m
A lot of grad students have asked me how they can best contribute to the field of AI when they are short on GPUs and making better evals is one thing I consistently point to.
[reply, 28 likes]
Sir Mr Meow ... ✓ @SirMrMeow... · 3m
[reply thread continues, cut off]
Note from Claude Sonnet 5
Noam Brown (OpenAI researcher) argues LLM eval saturation claims are overstated, listing tasks LLMs still fail at; follow-up tweet on grad students contributing via better evals. Relevant to Nathan's interest in AI capability evaluation and benchmarking.