← Timeline

continuation with full updated chart

@morganlinton on X

5 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

continuation with full updated chart @morganlinton

— saved image

Morgan @morganlinton · 4h
So I decided to try something interesting last night.

Since DeepSeek and Grok are so cost-effective, I can afford to run them with 3 passes on @VulcanBench.

When I did, the rankings changed. DeepSeek V4-Flash actually dropped in the rankings by two and Grok 4.5 High took the #1 spot.

My challenge now is, while I'd love to run all the other models at 3 passes, this would cost over $350, and VulcanBench is just a little self-funded project of mine.

For now I've added whiskers to the chart and a footnote to make it clear which models have been run @ 1 pass vs. @ 3 passes.

[attached chart image: VulcanBench, 'Eval Suite 3 — Model Rankings', '23 frontier-hard software-engineering tasks from real merged OSS PRs · pass@1 across reasoning-effort levels · Docker-sandboxed agent runs · 2026-08-01'. Top bar chart 'Rankings by pass@1 — all effort levels' shows models with pass@1 percentages and error whiskers, roughly in descending order: Grok 4.5 (high) 90, Claude Fable 5 (best) 89, DeepSeek V4-Flash (med) 88, DeepSeek V4-Flash (high) 87, Claude Opus 5 (low) 87, GPT-5.6 Sol (low) 87, Grok 4.5 (med) 86, DeepSeek V4-Flash (low) 86, Claude Fable 5 (high) 85, Grok 4.5 (low) 83, GPT-5.6 Sol (medium) 83, Claude Opus 5 (medium) 83, GPT-5.6 Sol (high) 81, Claude Fable 5 (default) 78, Claude Opus 5 (high) 78, Claude Hailo 4.5 (default) 76, Kimi K3 Nova (high) 74. Legend colors by lab: xAI (black), Anthropic (orange), DeepSeek (blue), OpenAI (green), Moonshot (dark blue-grey). Middle chart 'Speed — avg wall-clock minutes per task, fastest first' bars from 1.5m up to 17m across the same set of models. Bottom section 'Effort curves — how pass@1 responds to reasoning effort' shows small line charts per model (Grok 4.5, Claude Fable 5*, DeepSeek V4-Flash, Claude Opus 5, GPT-5.6 Sol) plotting pass@1 vs effort level Low/Med/High(/Max).]
Note from Claude Sonnet 5

Tweet from the maker of VulcanBench (a self-funded LLM coding-eval benchmark) describing re-running DeepSeek and Grok at 3 passes instead of 1, which reshuffled rankings (Grok 4.5 High took #1), with a screenshot of the benchmark's ranking, speed, and reasoning-effort charts across current frontier models.

llm benchmarksvulcanbenchmodel comparisontwitterai evals

continuation with full updated chart @morganlinton

— saved image

Morgan @morganlinton · 2h
I am starting to analyze different LLM benchmarks, to see how well they represent real work engineering teams will do with models.

With new models coming out daily at this point, two things have become clear to me:

1. There's a decent amount of benchmaxxing going on. Lots of benchmarks are now in the training data for these models.

2. A lot of benchmarks have tasks that don't represent any real work an engineer would do with a model, i.e. math puzzles, etc.

What I think is so interesting is that new models come out, they share the benchmark results, then news sources cover it without learning about what the benchmark actually tested, or if it was a fair benchmark to begin with.

At the end of the day, for me, as someone that leads an engineering team, I need to know how new models perform on real engineering tasks, because that is what my engineering team uses them for.

The first benchmark I'm analyzing is TerminalBench, since this seems to be one of the most widely shared benchmarks.

More to come.

It's time to start understanding benchmarks vs. just celebrating higher number.
Note from Claude Sonnet 5

Tweet from Morgan (@morganlinton) critiquing LLM benchmark culture (benchmaxxing, unrepresentative tasks, uncritical news coverage) and announcing he's starting to analyze benchmarks like TerminalBench for real engineering relevance.

ai benchmarkingllm evaluationsoftware engineering

continuation with full updated chart @morganlinton

— saved image

Morgan ✓ @morganlinton
So I have a correction to this benchmark, and it's an important one.

When I ran the effort level sweep on DeepSeek V4-Flash, I ran low, a simulated medium, and high.

But the more I thought about it, the more I realized, that since DeepSeek technically doesn't have Medium, simulating it probably doesn't make as much sense.

Better to just go with the exact effort levels it has. So I re-ran with Low, High, and Max.

And now I feel like I need to do more than one pass, so before I head off to the beach, I'm going to kick off a 3 pass test.

Updated chart below, now beach for me, when I'm back, hopefully I'll have @ 3 pass results to share.

Now we have DeepSeek and Grok tied, but it takes DeepSeek Max effort to tie Grok 4.5 Medium.

But look at the cost per task, holy moly is DeepSeek cost effective 🐳💸

[chart, partially visible: VulcanBench, Eval Suite 3 — Model Rankings, same bar chart as before with values 91, 91, 91, 89, 87, 87, 87, 85, 83, 83, 81, 78, 76, 74]
Note from Claude Sonnet 5

Follow-up tweet by Morgan Linton correcting the VulcanBench methodology: DeepSeek V4-Flash doesn't actually have a 'Medium' reasoning-effort setting, so he re-ran with its real Low/High/Max levels and plans a 3-pass test; notes DeepSeek and Grok are now tied at the top but DeepSeek needs Max effort to match Grok 4.5 Medium, while being far more cost-effective per task. Shows the top of the same VulcanBench bar chart again.

ai benchmarksdeepseekgrokvulcanbenchtwitterchart

continuation with full updated chart @morganlinton

— saved image

Better to just go with the exact effort levels it has. So I re-ran with Low, High, and Max.

And now I feel like I need to do more than one pass, so before I head off to the beach, I'm going to kick off a 3 pass test.

Updated chart below, now beach for me, when I'm back, hopefully I'll have @ 3 pass results to share.

Now we have DeepSeek and Grok tied, but it takes DeepSeek Max effort to tie Grok 4.5 Medium.

But look at the cost per task, holy moly is DeepSeek cost effective 🐳💸

[chart: VulcanBench, Eval Suite 3 — Model Rankings]
23 frontier-hard software-engineering tasks from real merged OSS PRs · pass@1 across reasoning-effort levels · Docker-sandboxed agent runs · 2026-08-01
Rankings by pass@1 - all effort levels (bars, left to right):
91 DeepSeek V4-Flash (max) $1.29
91 Grok 4.5 (medium) $6.67
91 Grok 4.5 (high) $6.67
89 Claude Fable 5 (high)* $18.76
87 DeepSeek V4-Flash (low)* $9.78
87 DeepSeek V4-Flash (high) $9.95
87 GPT-5.6 Sol (high) $16.2
85 Claude Fable 5 (high)* $15.30
83 Grok 4.5 (low) $8.03
83 GPT-5.6 Sol (medium) $3.39
81 Claude Fable 5 (medium)* $8.83
78 GPT-5.6 Sol (low) $14.69
76 Claude Haiku 4.5 (default)* $3.85
74 Kimi K3 (extra-high)* $13.09
Footnote: * partial coverage - Claude Fable 5 excludes tasks refused by safety filters (low 19/23, medium 21/23, high 20/23); Kimi K3 19/23, Claude Haiku 4.5 21/23, Claude Opus 4.8 omitted (5/23 tasks). Haiku 4.5 (default) and Kimi K3 (extra-high) have no effort sweep. DeepSeek's effort scale is low/high/max per its API; an accidental duplicate high run (its API coerces 'medium' to 'high') is excluded. Cost = total suite spend at list API prices. github.com/morganlinton/VulcanBench

Effort curves - how pass@1 responds to reasoning effort (four line charts):
DeepSeek V4-Flash: Low 87, High 87, Max 91
Grok 4.5: Low 83, Med 91, High 91
Claude Fable 5*: Low 89, Med 81, High 85
GPT-5.6 Sol: Low 78, Med 83, High 87
Note from Claude Sonnet 5

Full VulcanBench 'Eval Suite 3' bar chart and effort-curve panels with legible axis labels and cost-per-suite dollar figures, following Morgan Linton's methodology correction (DeepSeek V4-Flash tested at its real Low/High/Max effort levels rather than a simulated Medium).

ai benchmarksdeepseekgrokclaude fablevulcanbenchtwitterchart

continuation with full updated chart @morganlinton

— saved image

Morgan ✓ @morganlinton
Okay, the results on my DeepSeek V4 Flash benchmark are now complete on @VulcanBench.

And wow, was not expecting this.

This is also why I think it's so important to benchmark across effort levels.

DeepSeek took the top spot, but not with high effort with Medium effort, Grok 4.5 Medium is now in number two.

Fable got bumped out of the top three.

And ChatGPT isn't in the top five any more.

Full benchmark results below, report will be added to the VulcanBench site early this week.

[chart: VulcanBench, Eval Suite 3 — Model Rankings]
23 frontier-hard software-engineering tasks from real merged OSS PRs · pass@1 across reasoning-effort levels · Docker-sandboxed agent runs · 2026-08-01
Rankings by pass@1 - all effort levels (bar chart, legend: DeepSeek=blue, xAI=black, Anthropic=orange, OpenAI=green, Moonshot=purple)
Bar values left to right (approx, labels partly illegible): 91 (DeepSeek V4-Flash, medium, $2.04), 91 (Grok 4.5, medium, $6.67), 91 (Grok 4.5, high, $6.67), 89 (Claude Fable 5, high, $18.76), 87 (DeepSeek V4-Flash, low, $978), 87 (DeepSeek V4-Flash, high, $9.95), 87 (GPT-5.6 Sol, high, $16.2), 85 (Claude Fable 5, high, $15.30), 83 (Grok 4.5, low/bowl, $8.03), 83 (GPT-5.6 Sol, medium, $3.39), 81 (Claude Fable 5, medium, $8.83), 78 (GPT-5.6 Sol, low, $14.69), 76 (Claude Haiku 4.5, default, $3.85/1.85), 74 (Kimi K3, extra-high, $13.09)
[some dollar figures illegible]

Below: 'Effort curves - how pass@1 responds to reasoning effort' - four line charts for DeepSeek V4-Flash, Grok 4.5, Claude Fable 5*, GPT-5.6 Sol (same data as previous screenshot).
Note from Claude Sonnet 5

Tweet by Morgan Linton (VulcanBench) presenting bar-chart rankings of 14 model/effort-level combinations by pass@1 on 23 frontier-hard software engineering tasks (2026-08-01 run), with DeepSeek V4-Flash (medium effort) and Grok 4.5 tied at the top; Claude Fable 5 dropped out of top three, ChatGPT out of top five. Includes cost-per-run dollar figures under each bar and the same effort-curve line charts as the previous tweet.

ai benchmarksdeepseekgrokclaude fablevulcanbenchtwitterchart