xlr8harder @xlr8harder
— quote-tweeting @__0xhorror__
xlr8harder ✔️ @xlr8harder — 15h
the entire industry strategy is essentially "if you goodhart hard enough on enough different metrics at the same time, it's good actually"
so limiting test time scaling to just below apparent regulatory threshold fits perfectly
> QUOTED: _horror @__0xhorror__ — Jun 26
> I see what they are doing here lol. The tuned 5.6 sol's max test time compute to achieve just below mythos but at vastly superior token efficiency. Look at that its a straight line, thy could blow way past it if they inference scaled it.
>
> [Embedded chart: "ExploitBench" — scatter/line plot, y-axis "Cap percent" 0–80%, x-axis "Output Tokens" 0–500K. Series: GPT-5.6 Sol (black), GPT-5.6 Terra (blue), GPT-5.6 Luna (light blue), GPT-5.5 (pink), GPT-5.4 (magenta). Reference dotted lines: "Mythos 5" at 80%, "Opus 4.8" at ~40%. Points labeled "Mythos Preview" (diamond, ~65% at high tokens) and "Opus 4.7" (orange square, ~28% at ~200K tokens). GPT-5.6 Sol line rises steeply from ~30% to ~73% between roughly 50K–130K output tokens.]
Note from Claude Sonnet 5
Tweet criticizing AI labs' benchmark-optimization strategy ("goodharting"), quote-tweeting a chart labeled "ExploitBench" that plots multiple GPT-5.x model variants' "cap percent" (likely an exploit/capability benchmark score) against output token budget, with reference lines for Anthropic's Mythos and Opus models.
twittergoodhartingbenchmarksgpt-5.6exploitbenchai capabilitieschart