← Timeline

tom cunningham

@testingham on X

2 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

tom cunningham @testingham

reposted by METR

[repost icon] METR reposted tom cunningham @testingham — 2h Q: has AI accelerated aggregate discovery yet? My very general impresions, would love others' thoughts: 1. Vulnerability discovery is up a lot. Between 2X-5X increase in volume, although the average severity has fallen somewhat. 2. Algorithmic efficiency hasn't moved much. E.g. nanogpt, SAT solvers, compression efficiency, chess algorithmic ability. They haven't shown noticeable changes I believe. 3. Math is hard to judge. Erdős problems are certainly falling faster than historical rates, but it's mainly the obscure ones. We don't seem to be tightening upper and lower bounds on unknown quantities noticeably more quickly AFAICT. There are some prominent AI discoveries (e.g. unit distance) but I don't know what's the denominator, i.e. the average flow of comparably important results, my guess is it's still small. 4. Papers/code volume is up a lot. Total arXiv papers and github code is up a lot, but no good way of judging how the quality has changed.
Note from Claude Sonnet 5

Long-form text tweet (list format, numbered 1-4), no images. Cut off at item 4, likely continues beyond visible screenshot area.

ai-safetyr-and-d-automationmetrsingularity-trackingalgorithmic-progress

tom cunningham @testingham

reposted by Cheryl Wu

🔁 Cheryl Wu reposted tom cunningham @testingham · 7h My basic model of capabilities: LLMs are good at problems similar to those that appear in their training data. Training data largely reflects the world, and so LLMs are relatively good at problems that are common, relatively bad at problems that are rare. [Chart: "success" (y-axis) vs "common problems" → "rare problems" (x-axis). Three downward-sloping lines: "best human" (highest, shallowest slope), "avg human" (middle), "LLM" (blue, starts near best-human level on common problems but has the steepest slope, dropping below both human lines on rare problems, crossing avg human partway through]
Note from Claude Sonnet 5

A capabilities model argument (widely reposted) that LLM performance degrades faster than human performance as problems become rarer/more out-of-distribution, illustrated with a simple crossing-lines chart — LLMs start above average human but below best human on common problems, then fall below both on rare problems. Relevant to general AI capabilities/scaling discourse Nathan tracks (adjacent to the empirical singularity tracking and algorithmic-progress threads already in project memory).

llm capabilitiesscalingai researchtwittertom cunninghamout-of-distribution generalization