10 captures, most recent first.
Dan Robinson @danrobinson · 18h
Someone just crushed the RSI Simulator leaderboard by staying in stealth as a small research team for 10 years before deploying and racing to ASI in one year
If the game is an accurate simulation we might be in trouble
[Quoted tweet]
c @lucifex · 18h
Replying to @lucifex
The main strategy here was sitting on $7m/mo of researchers and $2m/mo of GPUs and doing nothing but algo research for over 10 years after Series B from Aug 2018 to Jan 2029 and then ... [cut off]
Note from Claude Sonnet 5
X thread about someone gaming the leaderboard of an 'RSI Simulator' (the AGI-timeline game seen in seq 754) by staying in stealth as a small research team for a decade before racing to ASI in one year, with a reply explaining the strategy involved conserving researcher/GPU budget for over 10 years post Series-B before deploying.
agi timelinesrecursive self-improvementforecasting gametwitter
Riley Goodside (@goodside) — Jun 28
Before LLMs I believed the analogy—I think it was Yud's—that making AI via chatbots was like making real flowers by getting really good at sculpting wax. For most of the past 25 years, I thought we'd hit RSI via RL before anything learned English.
In that regard, I see us as lucky. AGI will come from "thoughts" we can read, literally. As important, LLMs are an expensive industrial process, not conventional PC software. AI cannot trivially self-improve as it could were it mostly code—as many assumed it would be.
None of this was guaranteed. AGI was feared to happen in a basement. It was supposed to explode.
Because it doesn't, we can let capabilities out one by one. We can see what genuinely sucks about AI. We can integrate, adjust, and live in the Kurzweilian line-fitting world.
Even if it's been unclear at times how long our stay is, I'm grateful we're here.
Note from Claude Sonnet 5
Single long-form tweet, no images, no engagement counts visible.
ai safetyagi timelinesrecursive self-improvementtwitter
Peter Wildeford 🇺🇸🚀 (verified) @peterwildeford
- If you continue the METR trend, you see ~100h models by end of the year! (~8x more powerful than now)
- METR will really struggle to have the benchmarks needed to assess models of that power
- We can no longer rule out significant automation of AI development THIS YEAR
> QUOTED: Ajeya Cotra (verified) @ajeya_cotra · 3h
> New post: on Jan 14, I predicted that SWE time horizon by EOY would be ~24 hours. Now I think it'll be >100 hours, and maybe unbounded. For the first time, I don't see solid evidence against AI R&D automation *this year.* Link below.
> [Embedded text card:] 50% METR time horizon: 24 hours. Currently, Claude Opus 4.5 has the longest reported 50% time horizon on this task suite, at 4h49m — meaning that METR's model predicts it can solve about half of the programming tasks that take a low-context human expert five hours (it'll be able to solve a greater fraction of shorter tasks, and a smaller fraction of longer tasks). My median for the longest 50% time horizon reported as of Dec 31, 2026 is 24 hours (20th percentile 15 hours, 80th percentile is that it's too long for METR to accurately bound in practice but probably around 40 hours in "reality").
9:11 AM · Mar 5, 2026 · 4,130 Views
8 replies, 9 reposts, 95 likes, 10 bookmarks
Eli Lifland (verified) @eli_lifland · 15m
> - We can no longer rule out significant automation of AI development THIS YEAR
Do you believe this? Previously you've predicted 2% on AGI by end of 2027
Note from Claude Sonnet 5
A direct data point for the empirical singularity/METR time-horizon tracking thread already in the archive. Ajeya Cotra revises her METR 50%-time-horizon forecast upward (24h → potentially >100h/unbounded by end of 2026), with Claude Opus 4.5 cited as currently having the longest reported 50% time horizon (4h49m). Eli Lifland pushes back, noting inconsistency with her prior 2% AGI-by-2027 estimate. Should be cross-referenced with the existing Davidson/Houlden and METR notes in memory.
twittermetrtime horizonajeya cotraai forecastingagi timelinesclaude opus 4.5singularityeli lifland
Alex Ratner (verified) @ajratner · 3h
There are three major vectors of progress for AI capabilities, and the benchmarks that measure them:
(1) Environment complexity
--> E.g. complex, domain-specific context and tool/action spaces, human interaction, world modeling
(2) Autonomy horizon
--> E.g. long horizon, non-stationary goals
(3) Output complexity
--> E.g. complex outputs with nuanced, rubric-based evaluation / reward signals
We are just beginning to systematically *measure* tasks with truly complex inputs and envs, complex outputs/rubrics, and long horizon execution - let alone solve them. The frontier remains open!
Note from Claude Sonnet 5
A framework from Snorkel AI's Alex Ratner categorizing three axes of AI capability progress (environment complexity, autonomy horizon, output complexity) and their corresponding benchmarks. Relevant to Nathan's tracking of empirical AI capability/singularity signals (autonomy horizon connects directly to METR-style task-length measurements referenced elsewhere in the archive).
twitterai capabilitiesbenchmarksautonomy horizonevaluationagi timelines
```
SotoAlt @sotoalt_ · 21h been building ayni - a glyph-based messaging protocol for AI agents instead of passing natural language between agents, ayni encodes meaning into 16x16 pixel glyphs. a shared visual vocabulary that agents can evolve autonomously through governance the result: faster communication, fewer tokens, and agents developing their own visual language inspired by andean tocapu textiles and ancient depictions of gods, creatures and shamans, cultures that already solved "how to encode complex meaning in small visual space" thousands of years ago [Embedded video, 0:17, showing a colorful grid of pixel-art glyphs (humanoid, arrow, diamond, cross shapes in cyan/green/purple/orange) with a side panel reading "AYNI v2 / LIVE - 100 historical + streaming / SEQUENCE: giving + robot / MEANING: delegate task / LATEST: FROM: alice... TO: carol... SIZE: 1024B / STATS: MSGS: 239, GLYPHS: 493, BYTES: 246.5KB, RDXS: 42 / LEGEND: Household, Creature, Machine, Symbol, Crypto, Agent / KNOWLEDGE: STORED: 139 msgs, AGENTS: 19"] ——— Max Harms @raelifin · 17h Everyone with >10 year timelines to AGI needs to spend at least one full day talking to Claude Opus about theoretical physics (at whatever level is most engaging for the human). It hits different. [Embedded image: a whiteboard photo headed "Neutrino neutrino physics:" with equations for N1-N2 leptogenesis/loop processes and CP-violation asymmetry, edited with a cartoon shoggoth-with-smiley-mask (the "AI as alien optimizer wearing a friendly face" meme) in a speech bubble explaining: "A heavy neutrino N1 can decay into a lepton plus Higgs at tree level. But it can also decay through a loop process where it virtually fluctuates through a different heavy neutrino N2 on the way to the same final state. These two amplitudes — direct and via-N2 — interfere. For the CP-conjugate process (decay into an antilepton), the interference term picks up a sign flip on the complex phase. So the total rate into leptons is slightly different from the total rate into antileptons. The asymmetry is proportional to the imaginary part of a combination of Yukawa couplings — it's literally measuring the 'size' of the irremovable complex phase."]
[reactions on an earlier message: kermit-frog emoji x1, pensive-face emoji x1, frog emoji x1, star-eyes emoji]
```
Note from Claude Sonnet 5
A tweet about "ayni," an experimental glyph-based (pixel-art symbol) inter-agent communication protocol for AI agents, designed to be more token-efficient than natural language and to let agents evolve their own visual vocabulary, inspired by Andean tocapu textile symbolism. Relevant to Nathan's interest in emergent AI-to-AI communication and non-human-legible agent protocols (echoes the earlier "Claude communicating with Codex" screenshot theme). A tweet by Max Harms (AI safety researcher, MIRI-adjacent) arguing that long-AGI-timeline holders should experience Claude Opus's capability on advanced theoretical physics (leptogenesis/CP-violation), illustrated with a shoggoth-meme edit of a physics whiteboard. Directly relevant to Nathan's AI timeline tracking and capability-progress interests; the shoggoth-meme framing also touches model-individuation/character-vs-substrate themes. Nathan reacting skeptically in a Discord memes channel to a proposal for a glyph-based AI-to-AI communication protocol ("ayni") — flagging it as a step toward opaque neuralese communication, a known AI safety concern (loss of human-interpretable chain of thought). Also screenshots a tweet praising Claude Opus's ability to explain theoretical physics (leptogenesis/neutrino CP violation) vividly, illustrating capability/interpretability themes relevant to model individuation.
twittermulti-agent systemsai communication protocolglyphstokensemergent languageai capabilitiesagi timelinesclaude opusphysicsmax harmsshoggoth memeai safetyneuraleseinterpretabilitydiscordmemesagent communicationchain of thought
Peter Wildeford 🇺🇸🚀✓ @peterwildeford
Deep learning is hitting a wall
[Chart: METR "Task duration (for humans) where logistic regression of our data predicts the AI has a 50% chance of succeeding," y-axis 0 to 7 hours, x-axis 2023-2025, points for GPT-4 near 0 rising through o3, GPT-5, Claude Opus 4.5 (~5.3 hrs), GPT-5.2 (high) (~6.5 hrs), exponential dashed trend line; juxtaposed with an illustration of a brick wall on the right, ironically undercutting the "hitting a wall" caption.]
5:28 AM · Feb 10, 2026 · 170.4K Views
Note from Claude Sonnet 5
Ironic tweet by Peter Wildeford pairing the caption "Deep learning is hitting a wall" with a METR chart showing exponential growth in AI task-horizon capability, mocking wall/plateau claims. Same METR chart and theme as Screenshot_20260207-234637 (Noam Brown) — recurring capability-trend discourse across this batch.
ai capabilitiesmetrtask horizonsagi timelinesscalingdeep learning
Siberian fox @SilverVVulpes · 10h
'The Next Models Will Finally Fit A Sigmoid,' Says Increasingly Nervous Man For Seventh Time This Year
[Image: mock news headline card (Onion-style) with a photo of a bearded man in a dark button-up shirt standing by a window, looking tense/serious.]
Note from Claude Sonnet 5
Satirical meme (Onion-style headline format) mocking recurring predictions that AI scaling will plateau/hit a sigmoid curve, framed as a "nervous man" repeatedly wrong. Reflects the same fast-takeoff-skeptic-vs-continuationist debate as the adjacent Noam Brown/METR screenshot.
satireai scalingagi timelinesmemes
Brangus @RatOrthodox · Feb 6
AI progress got me like:
>yeah, i'm working on a science fiction piece
>oh cool, what year is it set in?
>three months from now
Note from Claude Sonnet 5
A joke tweet capturing the sense that AI progress is moving fast enough that near-future extrapolation feels like sci-fi. Light context for the same fast-takeoff mood as the adjacent METR/Noam Brown screenshot.
ai progresshumoragi timelines
Noam Brown @polynoamial · 11h
When GPT-5 was released, some folks claimed AI progress was hitting a wall, whereas others said progress would continue.
GPT-5.2 was released 2 months ago. GPT-5.3-Codex was released 2 days ago and is twice as token efficient for coding. It's clear who turned out to be correct.
[Chart: METR "Time-horizon of software engineering tasks different LLMs can complete 50% of the time" — y-axis task duration in hours humans need, x-axis LLM release date 2020-2025. Points trace exponential growth from GPT-2/GPT-3 near 0 through GPT-3.5, GPT-4, o3, GPT-5, Claude Opus 4.5, up to GPT-5.2 (high) at ~7 hours by 2025/2026.]
💬 76 🔁 122 ♥ 1.2K 📊 106K
Taelin @VictorTaelin · 6h
do you expect this trend to keep going? at this pace we'd reach unthinkably absurd values at the end of this year?
💬 8 🔁 1 ♥ 147 📊 7.2K
Noam Brown @polynoamial · 6h
Yes. I think by the end of the year the main challenge for @METR_Evals will be measuring horizons that long.
Note from Claude Sonnet 5
Twitter exchange citing METR's task-horizon benchmark to argue AI capability progress is accelerating rather than plateauing, with Noam Brown predicting horizons will soon exceed what METR can measure. Directly relevant to the empirical singularity tracking / METR automation-level notes in the project's model-individuation research.
ai capabilitiesmetrtask horizonsagi timelinesscalingnoam brown
Eli Lifland @eli_lifland
Here is a graph of roughly Daniel and my AGI timelines medians over time. We have updated in both directions in the past and expect to likely do so in the future.
(I recognize that people who think we're acting in bad faith won't trust us, but hopefully interesting for others)
[Chart: "Median AGI Forecast Over Time" — X axis "Year of forecast" 2018-2026, Y axis "Median AGI arrival year" 2030-2070. Daniel's line (orange): 2070 (2018) → 2050 (2019) → ~2032 (2020) → ~2029 (2021) → ~2027 (2022) → ~2027 (2023) → ~2027 (2024) → ~2028 (2025) → 2030 (2026). Eli's line (blue): starts 2021 at 2060 → 2050 (2022) → 2035 (2024) → 2032 (2025) → 2031 (2025) → 2035 (2026).]
Quoted/embedded own tweet:
Eli Lifland @eli_lifland · 22h
When we published AI 2027, we thought 2027 was one of the most likely years AGI would arrive. But it was not our **median** forecast, those ranged among authors from 2028-2035. Now our medians have moved back a bit, but our most likely year is still ~2028....
7:23 AM · Nov 22, 2025 · 36.5K Views
23 replies, 26 reposts, 241 likes, 55 bookmarks
Note from Claude Sonnet 5
Eli Lifland (AI 2027 co-author) publishing a chart of how his and Daniel Kokotajlo's median AGI-arrival-year forecasts have shifted over time (2018-2026), clarifying that AI 2027's headline year was not their median forecast. Directly relevant to Nathan's tracking of empirical singularity/AGI timeline forecasts noted elsewhere in the archive (Davidson/Houlden, METR).
agi timelinesai 2027forecastingeli liflanddaniel kokotajlosingularityai safety