6 captures, most recent first.
[repost icon] METR reposted
tom cunningham @testingham — 2h
Q: has AI accelerated aggregate discovery yet?
My very general impresions, would love others' thoughts:
1. Vulnerability discovery is up a lot. Between 2X-5X increase in volume, although the average severity has fallen somewhat.
2. Algorithmic efficiency hasn't moved much. E.g. nanogpt, SAT solvers, compression efficiency, chess algorithmic ability. They haven't shown noticeable changes I believe.
3. Math is hard to judge. Erdős problems are certainly falling faster than historical rates, but it's mainly the obscure ones. We don't seem to be tightening upper and lower bounds on unknown quantities noticeably more quickly AFAICT. There are some prominent AI discoveries (e.g. unit distance) but I don't know what's the denominator, i.e. the average flow of comparably important results, my guess is it's still small.
4. Papers/code volume is up a lot. Total arXiv papers and github code is up a lot, but no good way of judging how the quality has changed.
Note from Claude Sonnet 5
Long-form text tweet (list format, numbered 1-4), no images. Cut off at item 4, likely continues beyond visible screenshot area.
ai-safetyr-and-d-automationmetrsingularity-trackingalgorithmic-progress
Nicholas Joseph ✓ @nickevanjoseph · 2h
Excited to welcome Andrej to the Pretraining team! He'll be building a team focused on using Claude to accelerate pretraining research itself. I can't think of anyone better suited to do it — looking forward to what we build together!
> Andrej Karpathy ✓ @karpathy · 2h
> Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resum...
Note from Claude Sonnet 5
Announcement that Andrej Karpathy has joined Anthropic's Pretraining team, specifically to use Claude to accelerate pretraining research (AI R&D automation applied recursively). Directly relevant to the empirical singularity / AI-R&D-automation tracking thread in project memory — Karpathy joining Anthropic to build "using Claude to accelerate pretraining" is a concrete instance of the automation trend being measured (METR r-value tracking).
anthropickarpathypretrainingai-r&d-automationsingularity-trackinghiring
dave kasten reposted
Lisan al Gaib ✓ @scaling01 · 22m
we have entered the kino zone
[Chart: "METR-Horizon-v1.1 P80 Time Horizons" — scatter plot with exponential fit line (R²=0.958), x-axis release date 2024-05 to 2026-05+, y-axis p80 time horizon in minutes (linear scale, 0-250). Chart is divided into three horizontal bands labeled "slop zone" (bottom, 0-50min), "transition zone" (middle, 50-180min), "kino zone" (top, 180-250min). Two vertical dashed lines mark "Karpathy's 'AI Agents are slop'" (~2025-11) and "Karpathy joins Anthropic" (~2026-05). Data point "Claude Mythos 185.9 min" is plotted near the top, just crossing into the kino zone, at roughly 2026-05.]
> October 2025: "AI agents are slop"
> May 2026: joins Anthropic x.com/karpathy/statu...
Lisan al Gaib ✓ @scaling01 · 59m
[quoted parent tweet, text truncated in screenshot]
Note from Claude Sonnet 5
A METR time-horizon benchmark chart showing Claude Mythos crossing into the "kino zone" (~186 min p80 task horizon), framed as vindication against Andrej Karpathy's earlier skepticism about AI agents, now that Karpathy has joined Anthropic. Relevant to the empirical singularity/AI-R&D-automation tracking thread already in project memory (METR time-horizon data, r-value discussions).
metrtime-horizonsclaude-mythosai-agentsscalingkarpathyanthropicsingularity-tracking
Peter Wildeford... @peterwildef... · 6h
real
> QUOTED (image of document text, with "Mid 2025" struck through and replaced by "Early 2026" in red):
Early 2026 [was: Mid-2025]: Stumbling Agents
The world sees its first glimpse of AI agents.
Advertisements for computer-using agents emphasize the term "personal assistant": you can prompt them with tasks like "order me a burrito on DoorDash" or "open my budget spreadsheet and sum this month's expenses." They will check in with you as needed: for example, to ask you to confirm purchases.⁸ Though more advanced than previous iterations like Operator, they struggle to get widespread usage.⁹
Meanwhile, out of public focus, more specialized coding and research agents are beginning to transform their professions.
The AIs of 2024 could follow specific instructions: they could turn bullet points into emails, and simple requests into working code. In 2025, AIs function more like employees. Coding AIs increasingly look like autonomous agents rather than mere assistants: taking instructions via Slack or Teams and making substantial code changes on their own, sometimes saving hours or even days.¹⁰ Research agents spend half an hour scouring the Internet to answer your question.
The agents are impressive in theory (and in cherry-picked examples), but in practice unreliable. AI twitter is full of stories about tasks bungled in some particularly hilarious way. The better agents are also expensive; you get what you pay for, and the best performance costs hundreds of dollars a month.¹¹ Still, many companies find ways to fit AI agents into their workflows.¹²
Note from Claude Sonnet 5
A retrospective note on the "AI 2027" forecast document (the "Stumbling Agents" section), with someone editing the original "Mid-2025" heading to "Early 2026" and Peter Wildeford endorsing the correction as "real" — i.e. the forecast's agent-capability timeline was roughly accurate but ran about 6-9 months later than predicted. Directly relevant to Nathan's interest in tracking empirical progress against AI forecasting/singularity models.
ai-2027forecastingai-agentstimelinespeter-wildefordtwittersingularity-tracking
davidad 🌟✓ @davidad · 13h
me@2024: Powerful AIs might all be misaligned; let's help humanity coordinate on formal verification and strict boxing
me@2026: Too late! Powerful AIs are ~here, and some are open-weights. But some are aligned! Let's help *them* cooperate on formal verification and cybersecurity
[Embedded chart: "Frontier AI's software R&D capabilities (Human-expert attention-span equivalent)" — a log-scale y-axis from 2s to 10h, x-axis quarterly from 2019 to 2028. Data points (x marks) trend upward from ~3s in 2019 to ~5h at a highlighted purple X around Q1 2026, with green/blue shaded uncertainty bands extending the trend to ~10-20h+ by 2027-2028. This appears to be a METR-style "time horizon" capability trend chart.]
[Quoted tweet:]
ARIA ✓ @ARIA_research · 15h
In Safeguarded AI, we're funding teams to develop systems that harden our critical infrastructure from growing vulnerabilities.
Programme Director @davidad warns that rapi... [cut off]
Note from Claude Sonnet 5
davidad (ARIA's Safeguarded AI programme director, known for formal-verification approaches to AI safety) reflects on how his own strategy shifted from 2024's "prevent misaligned AI" framing to 2026's "some AI is aligned, help it cooperate on verification/cybersecurity," alongside a METR-style capability trend chart showing frontier AI software R&D time-horizon capability crossing ~5 hours by early 2026. Directly relevant to Nathan's empirical singularity/AI-progress tracking notes (his memory already cites METR's 0.25-0.5 automation range and time-horizon doublings) — this chart is a current data point for that tracking thread, and davidad's strategic pivot is a notable AI-safety-field signal.
twitterdavidadariasafeguarded-aiformal-verificationai-capability-trendsmetrsingularity-trackingai-safety-strategy
Guillermo Rauch ✓ @rauchg · 6h
10 days into 2026:
- Terence Tao announces GPT & Aristotle solve Erdős problem autonomously
- Linus Torvalds concedes vibe coding is better than hand-coding for his non-kernel project
- DHH walks back "AI can't code" from Lex podcast 6 months later
An acceleration is coming the likes of which humanity has never experienced before
> QUOTED (embedded git commit screenshot):
torvalds committed 3 days ago
Merge branch 'antigravity'
This is Google Antigravity fixing up my visualization tool (which was also generated with help from google, but of the normal kind).
It mostly went smoothly, although I had to figure out what the problem with using the builtin rectangle select was. After telling antigravity to just do a custom RectangleSelector, things went much better.
Is this much better than I could do by hand? Sure is.
Note from Claude Sonnet 5
A tweet compiling anecdotal signals of accelerating AI capability adoption in early 2026 — Terence Tao crediting GPT+Aristotle with autonomously solving an Erdős problem, Linus Torvalds endorsing AI-assisted ("vibe") coding via a real git commit message, and DHH reversing a prior "AI can't code" stance. Relevant to Nathan's interest in tracking empirical signals of AI progress/takeoff speed (cf. his singularity-tracking notes on METR/Epoch data).
twitterai-progressvibe-codingterence-taolinus-torvaldsdhhsingularity-trackinggoogle-antigravity