Håvard Ihle @htihle
Håvard Ihle @htihle · 13h
Claude opus 4.6 (adaptive) takes the lead on WeirdML with 77.9% ahead of gpt-5.2 (xhigh) at 72.2%.
It sets a new high score on 3 tasks including scoring 73% on the hardest task (digits_generalize) up from 59%.
Opus 4.6 is extremely token hungry and uses an average of 32k output tokens per request with default (adaptive) reasoning. Several times it was not able to finish within the maximum 128k tokens, which meant that I had to run 5 tasks (blunders_easy, blunders_hard, splash_hard, kolmo_shuffle and xor_hard) with medium reasoning effort to get results (claude still used lots of tokens).
Because of the high cost, opus 4.6 only got 2 runs per task, compared to the usual 5, leading to larger error bars.
[Chart 1: "WeirdML Results Summary" — bar chart of average accuracy across 17 tasks and cost per run (USD) for various models, top score 0.903, Claude Opus entries around 0.779/0.722/0.699 etc. with costs ranging ~$0.22–$7.65]
[Chart 2: "WeirdML: Interactive Model Comparison" scatter plot, average cost per run (USD, log scale) vs performance, showing models from OpenAI, Anthropic, Google, DeepSeek, xAI, Meta, Mistral, Qwen, Moonshot]
> QUOTED: Håvard Ihle @htihle · Jun 27, 2025
> WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other ...
Note from Claude Sonnet 5
Benchmark results from the WeirdML evaluation suite showing Claude Opus 4.6 taking the top score (77.9%) ahead of GPT-5.2, with a note on Opus 4.6's high token/cost usage. Relevant to tracking frontier model capability comparisons over time, adjacent to the project's interest in Claude-generation capability jumps.
benchmarksweirdmlclaude opus 4.6gpt-5.2llm evaluationai capabilities