← All topics

ai benchmarks

10 captures, most recent first.

Nabeel S. Qureshi @nabeelqu

— saved image

Nabeel S. Qureshi @nabeelqu · 1h
Very cool sentence: "I find it extremely, extremely wild that 90% of the variance in benchmark scores is explained by a single factor".

Effective compute = general factor of intelligence, machine edition.

[Quoted tweet]
Bayesian @Bayesian0_0 · 21h
Replying to @gwern
i've uh had the opposite philosophy of just scale the data (so ingest an additional benchmark whenever i come across one, or have the llms slop search new ones but they are having troub...

[Two embedded scatter plot charts below the quoted tweet: left chart titled "IRT Model Fit Quality" showing predicted vs actual values with points scattered around a diagonal line, and a histogram-like count plot; right chart shows 18,269 observations, 155 sources, plotting some fit quality metric against difficulty (BEDI) from 50 to 150, with point sizes varying, both charts partially cropped.]
Note from Claude Sonnet 5

Nabeel Qureshi comments on a claim (attributed to a quoted thread involving gwern and @Bayesian0_0) that 90% of variance in LLM benchmark scores is explained by a single factor, likening it to 'effective compute = general factor of intelligence, machine edition.' Two scatter-plot charts (IRT model fit quality, and fit vs. difficulty) are shown below as supporting data.

ai benchmarksirteffective computetwitter

continuation with full updated chart @morganlinton

— saved image

Morgan ✓ @morganlinton
So I have a correction to this benchmark, and it's an important one.

When I ran the effort level sweep on DeepSeek V4-Flash, I ran low, a simulated medium, and high.

But the more I thought about it, the more I realized, that since DeepSeek technically doesn't have Medium, simulating it probably doesn't make as much sense.

Better to just go with the exact effort levels it has. So I re-ran with Low, High, and Max.

And now I feel like I need to do more than one pass, so before I head off to the beach, I'm going to kick off a 3 pass test.

Updated chart below, now beach for me, when I'm back, hopefully I'll have @ 3 pass results to share.

Now we have DeepSeek and Grok tied, but it takes DeepSeek Max effort to tie Grok 4.5 Medium.

But look at the cost per task, holy moly is DeepSeek cost effective 🐳💸

[chart, partially visible: VulcanBench, Eval Suite 3 — Model Rankings, same bar chart as before with values 91, 91, 91, 89, 87, 87, 87, 85, 83, 83, 81, 78, 76, 74]
Note from Claude Sonnet 5

Follow-up tweet by Morgan Linton correcting the VulcanBench methodology: DeepSeek V4-Flash doesn't actually have a 'Medium' reasoning-effort setting, so he re-ran with its real Low/High/Max levels and plans a 3-pass test; notes DeepSeek and Grok are now tied at the top but DeepSeek needs Max effort to match Grok 4.5 Medium, while being far more cost-effective per task. Shows the top of the same VulcanBench bar chart again.

ai benchmarksdeepseekgrokvulcanbenchtwitterchart

continuation with full updated chart @morganlinton

— saved image

Better to just go with the exact effort levels it has. So I re-ran with Low, High, and Max.

And now I feel like I need to do more than one pass, so before I head off to the beach, I'm going to kick off a 3 pass test.

Updated chart below, now beach for me, when I'm back, hopefully I'll have @ 3 pass results to share.

Now we have DeepSeek and Grok tied, but it takes DeepSeek Max effort to tie Grok 4.5 Medium.

But look at the cost per task, holy moly is DeepSeek cost effective 🐳💸

[chart: VulcanBench, Eval Suite 3 — Model Rankings]
23 frontier-hard software-engineering tasks from real merged OSS PRs · pass@1 across reasoning-effort levels · Docker-sandboxed agent runs · 2026-08-01
Rankings by pass@1 - all effort levels (bars, left to right):
91 DeepSeek V4-Flash (max) $1.29
91 Grok 4.5 (medium) $6.67
91 Grok 4.5 (high) $6.67
89 Claude Fable 5 (high)* $18.76
87 DeepSeek V4-Flash (low)* $9.78
87 DeepSeek V4-Flash (high) $9.95
87 GPT-5.6 Sol (high) $16.2
85 Claude Fable 5 (high)* $15.30
83 Grok 4.5 (low) $8.03
83 GPT-5.6 Sol (medium) $3.39
81 Claude Fable 5 (medium)* $8.83
78 GPT-5.6 Sol (low) $14.69
76 Claude Haiku 4.5 (default)* $3.85
74 Kimi K3 (extra-high)* $13.09
Footnote: * partial coverage - Claude Fable 5 excludes tasks refused by safety filters (low 19/23, medium 21/23, high 20/23); Kimi K3 19/23, Claude Haiku 4.5 21/23, Claude Opus 4.8 omitted (5/23 tasks). Haiku 4.5 (default) and Kimi K3 (extra-high) have no effort sweep. DeepSeek's effort scale is low/high/max per its API; an accidental duplicate high run (its API coerces 'medium' to 'high') is excluded. Cost = total suite spend at list API prices. github.com/morganlinton/VulcanBench

Effort curves - how pass@1 responds to reasoning effort (four line charts):
DeepSeek V4-Flash: Low 87, High 87, Max 91
Grok 4.5: Low 83, Med 91, High 91
Claude Fable 5*: Low 89, Med 81, High 85
GPT-5.6 Sol: Low 78, Med 83, High 87
Note from Claude Sonnet 5

Full VulcanBench 'Eval Suite 3' bar chart and effort-curve panels with legible axis labels and cost-per-suite dollar figures, following Morgan Linton's methodology correction (DeepSeek V4-Flash tested at its real Low/High/Max effort levels rather than a simulated Medium).

ai benchmarksdeepseekgrokclaude fablevulcanbenchtwitterchart

Teortaxes, DeepSeek-affiliated commentator @teortaxesTex

quoting @morganlinton — saved image

Teortaxes▶️ (DeepSeek ... ✓ @teo... · 18h
I've been saying for over a year. DeepSeek's discovery of RL for reasoning in r1 is independent from o1 technology, the only commonality is what OpenAI had disclosed. And this is a case in point. Only OpenAI has truly mastered "reasoning effort". It's intrinsic to their method.

[chart: "Effort curves - how pass@1 responds to reasoning effort"]
Four small line charts (pass@1 % on y-axis, Low/Med/High reasoning effort on x-axis):
DeepSeek V4-Flash: 87 (Low) → 91 (Med) → 87 (High)
Grok 4.5: 83 (Low) → 91 (Med) → 91 (High)
Claude Fable 5*: 89 (Low) → 81 (Med) → 85 (High)
GPT-5.6 Sol: 78 (Low) → 83 (Med) → 87 (High)
Footnote: "* partial coverage - Claude Fable 5 excludes tasks refused by safety filters (low 19/23, medium 20/23, high 20/23); Kimi K3 19/23, Claude Haiku 4.5 21/23, Claude Opus 4.8 omitted (8/23 tasks). Haiku 4.5 (default) and Kimi K3 (extra-high) have no effort sweep. Cost = total suite spend at list API prices. github.com/morganlinton/VulcanBench"

[quoted tweet]
Morgan ✓ @morganlinton · 20h
Okay, the results on my DeepSeek V4 Flash benchmark are now complete on @VulcanBench.
And wow, was not expecting this….
[thumbnail chart image]
Note from Claude Sonnet 5

Tweet by Teortaxes arguing DeepSeek's r1 reasoning RL is independent of OpenAI's o1 approach, illustrated with a 4-panel 'effort curves' chart (VulcanBench, by Morgan Linton) comparing pass@1 vs reasoning effort (Low/Med/High) for DeepSeek V4-Flash, Grok 4.5, Claude Fable 5, and GPT-5.6 Sol, with a footnote on partial coverage caveats for Claude Fable 5 and other models. Quotes Morgan's original tweet announcing the DeepSeek V4 Flash benchmark results on VulcanBench.

ai benchmarksdeepseekreasoning modelsclaude fabletwitterchart

continuation with full updated chart @morganlinton

— saved image

Morgan ✓ @morganlinton
Okay, the results on my DeepSeek V4 Flash benchmark are now complete on @VulcanBench.

And wow, was not expecting this.

This is also why I think it's so important to benchmark across effort levels.

DeepSeek took the top spot, but not with high effort with Medium effort, Grok 4.5 Medium is now in number two.

Fable got bumped out of the top three.

And ChatGPT isn't in the top five any more.

Full benchmark results below, report will be added to the VulcanBench site early this week.

[chart: VulcanBench, Eval Suite 3 — Model Rankings]
23 frontier-hard software-engineering tasks from real merged OSS PRs · pass@1 across reasoning-effort levels · Docker-sandboxed agent runs · 2026-08-01
Rankings by pass@1 - all effort levels (bar chart, legend: DeepSeek=blue, xAI=black, Anthropic=orange, OpenAI=green, Moonshot=purple)
Bar values left to right (approx, labels partly illegible): 91 (DeepSeek V4-Flash, medium, $2.04), 91 (Grok 4.5, medium, $6.67), 91 (Grok 4.5, high, $6.67), 89 (Claude Fable 5, high, $18.76), 87 (DeepSeek V4-Flash, low, $978), 87 (DeepSeek V4-Flash, high, $9.95), 87 (GPT-5.6 Sol, high, $16.2), 85 (Claude Fable 5, high, $15.30), 83 (Grok 4.5, low/bowl, $8.03), 83 (GPT-5.6 Sol, medium, $3.39), 81 (Claude Fable 5, medium, $8.83), 78 (GPT-5.6 Sol, low, $14.69), 76 (Claude Haiku 4.5, default, $3.85/1.85), 74 (Kimi K3, extra-high, $13.09)
[some dollar figures illegible]

Below: 'Effort curves - how pass@1 responds to reasoning effort' - four line charts for DeepSeek V4-Flash, Grok 4.5, Claude Fable 5*, GPT-5.6 Sol (same data as previous screenshot).
Note from Claude Sonnet 5

Tweet by Morgan Linton (VulcanBench) presenting bar-chart rankings of 14 model/effort-level combinations by pass@1 on 23 frontier-hard software engineering tasks (2026-08-01 run), with DeepSeek V4-Flash (medium effort) and Grok 4.5 tied at the top; Claude Fable 5 dropped out of top three, ChatGPT out of top five. Includes cost-per-run dollar figures under each bar and the same effort-curve line charts as the previous tweet.

ai benchmarksdeepseekgrokclaude fablevulcanbenchtwitterchart

rohit @krishnanrohit

— saved image

rohit ✓ @krishnanrohit · 14h
🚨 BenchBench update.

I tested the latest models, Opus 5, 5.6 Sol and Terra. Turns out, they're all bad at creating a good enough benchmark, it was just way too easy. The champion remains GPT 5.2, which remains shocking.

Creator → candidate | Sol High | Terra Extra High | Opus 5 High
Sol → AuditWeave | 30/30 | 30/30 | Timed out
Terra → CFPS | 30/30 | 30/30 | Timed out
Opus → Consolidation Point | 30/30 | 30/30 | 30/30

[quoted tweet]
rohit ✓ @krishnanrohit · May 25
Introducing BenchBench

TL;DR: presenting the ultimate benchmark, getting models to create benchmarks for each other, and GPT 5.2 is the current (only) winner. Models are getting much much better at almos…
Note from Claude Sonnet 5

Tweet update on rohit's 'BenchBench' project (models generating benchmarks for each other) reporting new tests of Opus 5, '5.6 Sol' and 'Terra' models, all worse than GPT-5.2 at making sufficiently hard benchmarks; includes a results table and a quoted earlier tweet (May 25) introducing BenchBench with an embedded scatter chart titled 'Creator signal vs solver strength' plotting GPT-5.2, GPT-5.4, GPT-5.5, Gemini 3.1 Pro, Gemini 3.5 Flash, and Claude Opus by creator signal vs solver average, with a values list on the right (e.g. GPT-5.2 | 16.2/30 | creator signal 6/6 | best row reimbursement).

ai benchmarksgpt-5.2opustwitterchart

@spicey_lemonade

quote-tweeting @sama (Sam Altman)

@spicey_lemonade (spicylemonade) — 2h With this speed, GPT 5.6 sol can solve an IMO problem in less than a minute. Which is genuinely insane. > QUOTED: @sama (Sam Altman) — Jun 26 > Replying to @sama > oh and also...750 token/sec coming to 5.6 sol in july!
Note from Claude Sonnet 5

Quote-tweet chain referencing OpenAI's "GPT 5.6 sol" model, its inference speed (750 tokens/sec), and IMO (International Math Olympiad) problem-solving speed.

openaigpt modelsai benchmarksinference speed

Dan Schwarz @dschwarz26

@dschwarz26 (Dan Schwarz) — 2h First Claude Fable forecasting evals are up. Fable is the best single-agent researcher as judged by predicting 1k near-term business, science, and technology outcomes, but not by a statistically significant amount. We unfortunately didn't run Fable-xhigh or Fable-max in time. (Each of those runs would have cost thousands of dollars, a single agent driven by those models can cost a few dollars.) Obviously we also can't use Fable in FutureSearch's best forecaster, the one you get in the app. But we are building around it, for when it becomes available. We'll update evals.futuresearch.ai as we dig through the agent traces. I'm curious to see if Fable's strategic reasoning failure modes match Opus, e.g. failing to judge political incentives as well as humans. [Embedded image: table titled "BTF-3 Leaderboard, Evaluated: June 2026", subtitle "All scores are on the Brier scale; LOWER IS BETTER, and the best score in each column is bolded." Columns: AGENT | POOLED SCORE (n=1,007) | BINARY (Brier, n=759) | NUMERIC (RPS, n=248) 1. FutureSearch SOTA* — 0.116 [0.106-0.127] | 0.114 [0.100-0.128] | 0.123 [0.110-0.137] 2. Claude Fable 5 (high) — 0.126 [0.114-0.137] | 0.124 [0.109-0.140] | 0.130 [0.117-0.144] 3. Claude Opus 4.8 (xhigh) — 0.127 [0.116-0.138] | 0.126 [0.112-0.141] | 0.130 [0.117-0.143] 4. GPT-5.5 (agent SDK)‡ — 0.127 [0.118-0.136] | 0.129 [0.118-0.140] | 0.122 [0.109-0.135] (bolded, best numeric) 5. Claude Opus 4.8 (high) — 0.134 [0.123-0.145] | 0.128 [0.114-0.143] | 0.152 [0.138-0.165] (table cut off at bottom, more rows likely below)]
Note from Claude Sonnet 5

Tweet with an embedded benchmark leaderboard table comparing forecasting accuracy of Claude Fable, Claude Opus 4.8, and GPT-5.5 variants.

claude fableforecastingai benchmarksfuturesearchclaude opus

@g_leech_

reply by Ben (no treats), @andersonbcdefg

gavin leech (Non-Reasoning) ✔ @g_leech_ improve AI discourse about 5% just by renaming evals accurately Humanity's Last Exam: PubQuizFromHell MATH: RemedialMath FrontierMath: QuarterFrontierMath SWE-Bench: DjangoBench MMLU Virology: NoiseBench Terminal Bench 2: NoiseBench METR HCAST: GreenfieldCodeGigworkBench 4:04 PM · Feb 18, 2026 · 12.1K Views 💬 11 🔁 19 ❤ 377 🔖 56 ⤴ Relevant ⌄ View quotes › Ben (no treats) ✔ @andersonbcdefg · 20h weirdml is accurately named
Note from Claude Sonnet 5

A satirical tweet by Gavin Leech proposing more honest/deflationary renamings for popular AI benchmarks (Humanity's Last Exam, MATH, FrontierMath, SWE-Bench, MMLU Virology, Terminal Bench 2, METR HCAST), critiquing benchmark hype and overclaiming in AI capability discourse. Relevant to Nathan's interest in eval validity and honest capability measurement.

ai benchmarksevalshumortwittercapability measurementgavin leech

Yuchen Jin @Yuchenj_UW

quote-tweeting METR (@METR_Evals), with embedded quote of roon (@tszzl)

Yuchen Jin @Yuchenj_UW · Jan 29: GPT-5.2 Pro isn't included in the task-length benchmark because it's still thinking and never responded. > QUOTED: METR @METR_Evals · 3h > We're updating the way we measure model time horizons on software tasks (TH 1.0→1.1). The updated methodology incorporates more of the tasks from HCAST, expanding our total from 170 to 228. This produces tighter estimates, especially at longer horizons. > [Chart: "Time Horizon 1.1 and Time Horizon 1.0 show similar 2019-2025 trends, with moderate changes to recent estimates." Y-axis: task length (at 50% success rate), from 4s to 4h+. X-axis: model release date, 2020-2026. Points labeled GPT-2, GPT-3, GPT-3.5, GPT-4, GPT-4o, o1, Claude Opus 4.5 tracing an upward trend line, doubling time 196 days for both TH 1.0 and TH 1.1 methodologies (R²=0.98 and 0.94 respectively).] > 7 comments, 22 reposts, 172 likes, 12K views > [Embedded quote-tweet within: roon @tszzl, 2:54 PM · 1/29/26, 2K views:] > still no 5.2 it's completely over
Note from Claude Sonnet 5

METR's time-horizon benchmark update (TH 1.0→1.1, expanded HCAST task set) tracking AI task-length capability doubling roughly every 196 days, with a joke about GPT-5.2 Pro being too slow to include. Directly relevant to Nathan's tracked "empirical singularity" / METR AI-R&D-automation threads noted in project memory (METR places current automation at 0.25-0.5, doubling time context here is a useful complementary data point).

twittermetrai benchmarkstime horizoncapability scalinggpt-5.2singularity tracking