← All topics

llm benchmarks

4 captures, most recent first.

continuation with full updated chart @morganlinton

— saved image

Morgan @morganlinton · 4h
So I decided to try something interesting last night.

Since DeepSeek and Grok are so cost-effective, I can afford to run them with 3 passes on @VulcanBench.

When I did, the rankings changed. DeepSeek V4-Flash actually dropped in the rankings by two and Grok 4.5 High took the #1 spot.

My challenge now is, while I'd love to run all the other models at 3 passes, this would cost over $350, and VulcanBench is just a little self-funded project of mine.

For now I've added whiskers to the chart and a footnote to make it clear which models have been run @ 1 pass vs. @ 3 passes.

[attached chart image: VulcanBench, 'Eval Suite 3 — Model Rankings', '23 frontier-hard software-engineering tasks from real merged OSS PRs · pass@1 across reasoning-effort levels · Docker-sandboxed agent runs · 2026-08-01'. Top bar chart 'Rankings by pass@1 — all effort levels' shows models with pass@1 percentages and error whiskers, roughly in descending order: Grok 4.5 (high) 90, Claude Fable 5 (best) 89, DeepSeek V4-Flash (med) 88, DeepSeek V4-Flash (high) 87, Claude Opus 5 (low) 87, GPT-5.6 Sol (low) 87, Grok 4.5 (med) 86, DeepSeek V4-Flash (low) 86, Claude Fable 5 (high) 85, Grok 4.5 (low) 83, GPT-5.6 Sol (medium) 83, Claude Opus 5 (medium) 83, GPT-5.6 Sol (high) 81, Claude Fable 5 (default) 78, Claude Opus 5 (high) 78, Claude Hailo 4.5 (default) 76, Kimi K3 Nova (high) 74. Legend colors by lab: xAI (black), Anthropic (orange), DeepSeek (blue), OpenAI (green), Moonshot (dark blue-grey). Middle chart 'Speed — avg wall-clock minutes per task, fastest first' bars from 1.5m up to 17m across the same set of models. Bottom section 'Effort curves — how pass@1 responds to reasoning effort' shows small line charts per model (Grok 4.5, Claude Fable 5*, DeepSeek V4-Flash, Claude Opus 5, GPT-5.6 Sol) plotting pass@1 vs effort level Low/Med/High(/Max).]
Note from Claude Sonnet 5

Tweet from the maker of VulcanBench (a self-funded LLM coding-eval benchmark) describing re-running DeepSeek and Grok at 3 passes instead of 1, which reshuffled rankings (Grok 4.5 High took #1), with a screenshot of the benchmark's ranking, speed, and reasoning-effort charts across current frontier models.

llm benchmarksvulcanbenchmodel comparisontwitterai evals

xjdr @_xjdr

reposted by Sichu Lu

↻ Sichu Lu reposted xjdr ✓ @_xjdr i saw Terrence Tao use sol med to answer a lot of very complex problems in one of his chat logs. i became curious. i had a particularly sticky problem that was in my 'ai cant do this yet' pile that i was only very recently able to get sol ultra to solve correctly (the problem itself doesn't matter, its was a distributed systems problem blah blah) but i decided to see if i could get any other current modern model to solve it with enough prompt engineering and care (and time etc, etc). turns out, with enough careful prompting and patience i was able to get sol high, opus 5, k3, glm 5.2, gemini flash 3.6, muse 1.1 and grok 4.5 to solve it almost identically to the original solution without any hints. the only difference was the care i took in specifying the problem and a tiny bit of guidance along the way (no context hints ever) . i tried a lot of others but these were the only ones that were successful. however, i was fairly surprised that these _were_ successful. this sounds anecdotal but this has really impacted the way i think about which models i use and for what purpose and how much thought and effort should go into each prompt Last edited 9:04 PM · Jul 26, 2026 · 49K Views
Note from Claude Sonnet 5

Dark-mode X screenshot, plain text thread comparing frontier model problem-solving ability across many named models (Sol Ultra/High, Opus 5, K3, GLM 5.2, Gemini Flash 3.6, Muse 1.1, Grok 4.5).

llm benchmarksai capabilitiesmodel comparisontwitter

Adam Karvonen @a_karvonen

Adam Karvonen ✔ @a_karvonen Interesting snippet from the Qwen-3.5 blog post: "Compared to the Qwen3 series, the post-training performance gains in Qwen3.5 primarily stem from our extensive scaling of virtually all RL tasks and environments we could conceive." [Embedded chart: "Average Ranking vs. Environment Scaling" — line chart, x-axis "Number of Training Environments" (0 to 17500), y-axis "Average Ranking (lower is better)" (4 to 14, inverted). Two lines for Qwen3.5-397B-A17B Thinking (blue) and Non-Thinking (red), both rising steeply from ~Qwen3-Max baseline (bottom left, ~13-14 ranking, 0 environments) up to top right at ~15000 environments where Thinking reaches rank ~3-4 and Non-Thinking ~5. Reference points along the y-axis at 0 environments show static rankings for other models: Claude-Opus 4.5 Thinking (~2-3), Claude-Opus 4.5 (~4-5), Gemini-3 Pro (~5), GPT-5.2 High (~5-6), DeepSeek-V3.2-Thinking (~8-9), Kimi-K2.5 Thinking (~9-10), Qwen3-Max-Thinking (~10), Qwen3-Max (~13-14). Chart footnote: overall performance calculated by averaging ranking on BFCL-V4, VITA-Bench, DeepPlanning, Tool-Decathlon, and MCP-Mark.] 9:18 PM · Feb 16, 2026 · 5,676 Views
Note from Claude Sonnet 5

A tweet highlighting a chart from Qwen-3.5's release blog post showing that post-training gains scale strongly with the number of distinct RL training environments used, with Qwen3.5-397B-A17B Thinking approaching Claude Opus 4.5 Thinking on agentic/tool-use benchmarks as environment count scales up. Relevant to Nathan's tracking of capability scaling trends and RL environment scaling as an algorithmic-progress driver.

llm benchmarksqwenreinforcement learningscaling lawscapabilitiestwitterclaude opus

will brown @willccbb

quoting teej (@teej_m); replies from Ariel (@redtachyon) and Randall Bennett (@randallb)

will brown @willccbb how do you write evals for writing evals? > QUOTED: » teej @teej_m · 10h > How do you write evals for writing? x.com/matsonj/status... 5:00 PM · Jul 12, 2025 · 5,513 Views 💬3 🔁2 ♥46 🔖5 🔗 Ariel @redtachyon · 10h Who evaluates the evals? 💬 🔁 ♥1 📊130 🔗 Randall Bennett @randallb · 10h i have an answer... calibrating graders. 💬 🔁 ♥1 📊88 🔗
Note from Claude Sonnet 5

Twitter banter among AI/ML practitioners on the recursive problem of writing evaluations for evaluations (meta-evals), a real methodological issue in LLM benchmark design — tangentially relevant to AI safety eval work Nathan follows.

twitterevalsai researchllm benchmarksmeta-evaluation