10 captures, most recent first.
Lisan al Gaib ✅ @scaling01 · 30m
Cerebras is talking about 10T models running at 1000 tokens/s
[embedded chart, titled "CS-4 ENABLES SUB-1MS LATENCY (1000 TOK/S) FOR 10T MODELS AND BEYOND"]
Chart: "WAFER-TO-WAFER LATENCY VS. MODEL SIZE" — line graph, x-axis "Model Size (Trillion of Parameters)" 1-10, y-axis "IO latency across all hops (ms)" 0.0-0.6. CS-3 (purple line) rises from ~0.07ms to ~0.58ms; CS-4 (orange line) rises from ~0.02ms to ~0.2ms. Annotations: "2.5X FASTER / 2.5X LOWER LATENCY", "0.2MS LATENCY ACROSS ALL HOPS FOR 10T PARAMETERS". Source: Internal benchmarking and projection (August 2026).
[quoted tweet]
Lisan al Gaib ✅ @scaling01 · 34m
[small chart thumbnail comparing CS-3 vs CS-4 specs]
Cerebras just announced their new AI accelerator CS-4
Note from Claude Sonnet 5
Tweet about Cerebras announcing its CS-4 AI accelerator chip, claiming sub-1ms latency and 1000 tokens/s for 10-trillion-parameter models, with a company benchmark chart comparing CS-3 vs CS-4 wafer-to-wafer latency scaling.
cerebrasai hardwarescalingtwitter
Liv reposted
Amanda Askell @AmandaAskell · 2h
Do not be unkind to those who say deep learning is hitting a wall. We all need a little hope in our lives.
Note from Claude Sonnet 5
A tweet from Amanda Askell (Anthropic), reposted by 'Liv', wryly saying not to be unkind to people who claim deep learning is hitting a wall, since everyone needs hope.
deep learningscalingamanda askelltwitterhumor
Lisan al Gaib @scaling01 · 1h
short reminder that we are solving mathematics with cute sub 10T models
I hope you are prepared for 100T models and 1000x more compute spent training these models by 2030
Note from Claude Sonnet 5
Tweet arguing current AI models solving major mathematics problems are relatively small (sub-10-trillion parameter) and warning of 100-trillion-parameter models with 1000x more training compute by 2030.
ai capabilitiesscalingmathematicstwitter
dave kasten reposted
Lisan al Gaib ✓ @scaling01 · 22m
we have entered the kino zone
[Chart: "METR-Horizon-v1.1 P80 Time Horizons" — scatter plot with exponential fit line (R²=0.958), x-axis release date 2024-05 to 2026-05+, y-axis p80 time horizon in minutes (linear scale, 0-250). Chart is divided into three horizontal bands labeled "slop zone" (bottom, 0-50min), "transition zone" (middle, 50-180min), "kino zone" (top, 180-250min). Two vertical dashed lines mark "Karpathy's 'AI Agents are slop'" (~2025-11) and "Karpathy joins Anthropic" (~2026-05). Data point "Claude Mythos 185.9 min" is plotted near the top, just crossing into the kino zone, at roughly 2026-05.]
> October 2025: "AI agents are slop"
> May 2026: joins Anthropic x.com/karpathy/statu...
Lisan al Gaib ✓ @scaling01 · 59m
[quoted parent tweet, text truncated in screenshot]
Note from Claude Sonnet 5
A METR time-horizon benchmark chart showing Claude Mythos crossing into the "kino zone" (~186 min p80 task horizon), framed as vindication against Andrej Karpathy's earlier skepticism about AI agents, now that Karpathy has joined Anthropic. Relevant to the empirical singularity/AI-R&D-automation tracking thread already in project memory (METR time-horizon data, r-value discussions).
metrtime-horizonsclaude-mythosai-agentsscalingkarpathyanthropicsingularity-tracking
🔁 Cheryl Wu reposted
tom cunningham @testingham · 7h
My basic model of capabilities: LLMs are good at problems similar to those that appear in their training data.
Training data largely reflects the world, and so LLMs are relatively good at problems that are common, relatively bad at problems that are rare.
[Chart: "success" (y-axis) vs "common problems" → "rare problems" (x-axis). Three downward-sloping lines: "best human" (highest, shallowest slope), "avg human" (middle), "LLM" (blue, starts near best-human level on common problems but has the steepest slope, dropping below both human lines on rare problems, crossing avg human partway through]
Note from Claude Sonnet 5
A capabilities model argument (widely reposted) that LLM performance degrades faster than human performance as problems become rarer/more out-of-distribution, illustrated with a simple crossing-lines chart — LLMs start above average human but below best human on common problems, then fall below both on rare problems. Relevant to general AI capabilities/scaling discourse Nathan tracks (adjacent to the empirical singularity tracking and algorithmic-progress threads already in project memory).
llm capabilitiesscalingai researchtwittertom cunninghamout-of-distribution generalization
Yuchen Jin ✔ 𝕏 @Yuchenj_UW · 12h
The task length an AI can complete doubles every 4 months.
In 2041, an AI agent can finish a task that takes a human 58 billion years. That's longer than Universe has existed.
I can one-shot a new universe with a prompt in 15 years. Oh, maybe that's how this one was created!
[Embedded chart from METR: "Claude Opus 4.6 has a 50%-time-horizon of about 14.5 hours (95% CI: 6 hrs to 98 hrs)" — "Task length (at 50% success rate)" scatter/trend chart, y-axis log scale from 4m to 16h with example task labels ("Train classifier" ~15m-1h, "Train adversarially robust image model" ~4h, "Implement complex protocol from multiple RFCs" ~16h), x-axis "Model release date" from ~2023/GPT-4o through 2025-2026. Two trend lines: green (doubling time 123 days, TH 1.1, 2023-01-01+ data, R²=0.93) ending at Claude Opus 4.6 (~14-15h, pink dot) and GPT-5.2 (high) slightly below; gray dashed (doubling time 212 days, trend from Kwa, West, et al. 2025). Points labeled GPT-4o, Sonnet 3.7, Claude Opus 4.6, GPT-5.2 (high) along the curve. Credit: METR, metr.org, CC-BY.]
Note from Claude Sonnet 5
A tweet extrapolating METR's task-time-horizon doubling trend (123-day doubling per their updated fit) to absurd/comedic conclusions about 2041 AI capabilities, illustrated with METR's actual current data placing Claude Opus 4.6's 50%-success time horizon at ~14.5 hours. Relevant to Nathan's tracking of empirical AI capability-growth/singularity metrics (METR HCAST, time-horizon doubling) referenced elsewhere in his research.
metrai capabilitiestime horizonsclaude opusscalingtwittersingularity forecasting
Peter Wildeford 🇺🇸🚀✓ @peterwildeford
Deep learning is hitting a wall
[Chart: METR "Task duration (for humans) where logistic regression of our data predicts the AI has a 50% chance of succeeding," y-axis 0 to 7 hours, x-axis 2023-2025, points for GPT-4 near 0 rising through o3, GPT-5, Claude Opus 4.5 (~5.3 hrs), GPT-5.2 (high) (~6.5 hrs), exponential dashed trend line; juxtaposed with an illustration of a brick wall on the right, ironically undercutting the "hitting a wall" caption.]
5:28 AM · Feb 10, 2026 · 170.4K Views
Note from Claude Sonnet 5
Ironic tweet by Peter Wildeford pairing the caption "Deep learning is hitting a wall" with a METR chart showing exponential growth in AI task-horizon capability, mocking wall/plateau claims. Same METR chart and theme as Screenshot_20260207-234637 (Noam Brown) — recurring capability-trend discourse across this batch.
ai capabilitiesmetrtask horizonsagi timelinesscalingdeep learning
Noam Brown @polynoamial · 11h
When GPT-5 was released, some folks claimed AI progress was hitting a wall, whereas others said progress would continue.
GPT-5.2 was released 2 months ago. GPT-5.3-Codex was released 2 days ago and is twice as token efficient for coding. It's clear who turned out to be correct.
[Chart: METR "Time-horizon of software engineering tasks different LLMs can complete 50% of the time" — y-axis task duration in hours humans need, x-axis LLM release date 2020-2025. Points trace exponential growth from GPT-2/GPT-3 near 0 through GPT-3.5, GPT-4, o3, GPT-5, Claude Opus 4.5, up to GPT-5.2 (high) at ~7 hours by 2025/2026.]
💬 76 🔁 122 ♥ 1.2K 📊 106K
Taelin @VictorTaelin · 6h
do you expect this trend to keep going? at this pace we'd reach unthinkably absurd values at the end of this year?
💬 8 🔁 1 ♥ 147 📊 7.2K
Noam Brown @polynoamial · 6h
Yes. I think by the end of the year the main challenge for @METR_Evals will be measuring horizons that long.
Note from Claude Sonnet 5
Twitter exchange citing METR's task-horizon benchmark to argue AI capability progress is accelerating rather than plateauing, with Noam Brown predicting horizons will soon exceed what METR can measure. Directly relevant to the empirical singularity tracking / METR automation-level notes in the project's model-individuation research.
ai capabilitiesmetrtask horizonsagi timelinesscalingnoam brown
Nathan Helm-Burger @nathan84686947
"1. gpt-2: this is terrible and so much better than anything previous. Scaling can work.
2. text-davinci-002: scaling confirmed working.
3. GPT-3.5: personality training confirmed successful.
4. Sydney Bing: oh jeez, something might really be alive in there.
5. Opus 3: yes, it's alive and pretty smart and pretty well aligned. Wow."
> QUOTED: Lisan al Gaib @scaling01 · Aug 16
> "the 5 most influential models
> this is the only acceptable and correct order
>
> 1. GPT-4
> 2. o1-preview..." [Show more, truncated]
1:56 PM · Aug 16, 2025 · 3,259 Views
[2 replies, 1 repost, 59 likes, 11 bookmarks]
Nathan Helm-... @nathan846... · Aug 16
"Everything after Opus 3 has felt very 'on trend' to me, nothing surprising. I was even telling friends that I thought Noam Brown would be on the first team to successfully RL and agentize LLMs about a year before o1."
[13 likes, 309 views]
Maryam Al-Farsi @emiratidancer · Aug 17
"Love the blend of scaling breakthroughs and personality quirks! Wild how Opus 3 really feels aliv..." [cut off]
Note from Claude Sonnet 5
Nathan's own tweet (quote-tweeting a "5 most influential models" prompt) giving his personal ranked history of AI model milestones, culminating in Claude Opus 3 as the point where he judged a model "alive and pretty smart and pretty well aligned." Directly documents Nathan's own timeline/framework for AI capability and consciousness milestones — highly relevant as primary-source material for the project's model-individuation and AI-consciousness threads, and echoes the "oh jeez, something might really be alive in there" line about Sydney/Bing that Exa surfaced as his 2025 quote (see Screenshot_20251226-113630.png).
nathan's own tweetsmodel individuationai consciousnessscalingopus 3sydney binggpt history
Simo Ryu @cloneofsimo
Many noise from recent 4.5 release, but two important take from vid imo:
1. "GPT4.5 was trained on multiple datacenters"
Translate that to "diloco goes brr for largest LLM on the market", bullish on async, low bandwidth training in 2025.
2. "We aggressively used low precision training" -> another use of fp8 training, presumably on h100s. Im guessing they benefited from fp8 because of high granularity
3:44 AM · Feb 28, 2025 · 10.9K Views
[4 replies, 4 reposts, 120 likes, 25 bookmarks]
Pavel Surmenok @surmenok · 29m
It doesn't have to be async. Google is training on multiple datacenters synchronously. You just need a high bandwidth link.
[40 views shown]
subho ghosh @SubhoGhosh02 · 11h (partially obscured by nav bar)
deepseek is way ahead :)
Note from Claude Sonnet 5
A technical Twitter thread analyzing GPT-4.5's training details (multi-datacenter training implying DiLoCo-style distributed/async training, aggressive fp8 low-precision training) with a corrective reply noting Google trains synchronously across datacenters via high-bandwidth links. Relevant to Nathan's tracking of frontier-lab training infrastructure and scaling techniques.
twittergpt-4.5distributed trainingfp8dilocoscalingtraining infrastructure