— quoting Anthropic @AnthropicAI
gian @giansegato
i've been running and curating our tbench numbers for a few months (such an elegant eval btw, huge fan), and at some point late last year i got super nerdsniped when i noticed that slight changes in how memory is allocated to the pods running the tasks would have a surprising impact on the final scores
it sent me down a rabbit hole. as it turns out, what we consider objective measurements are actually more subjective than they seem, and the scaffold _around_ agentic coding evals is an active element affecting the final numbers, if not properly controlled for (which is harder than it seems)
just published the findings on the engineering blog today!
it's a niche topic, but to me it's an extremely interesting problem space. as we race to a world where models are smarter than most humans, having the right epistemological tools to deeply understand and measure their capabilities with high confidence is set to become ever harder, and thus ever more important to get right
> QUOTED: Anthropic ✅ @AnthropicAI · 9h
> New on the Engineering Blog: Quantifying infrastructure noise in agentic coding evals.
> Infrastructure configuration can swing agentic coding benchmarks by several percentage points—sometime...
Note from Claude Sonnet 5
An Anthropic engineer's thread on discovering that infrastructure/scaffold configuration (e.g. memory allocation to pods) introduces significant noise into agentic coding benchmark scores, framed as an epistemological problem for measuring model capability as models approach/exceed human level. Relevant to Nathan's interest in eval robustness and the reliability of capability measurements underlying singularity-timeline forecasts.
twitteranthropicevaluationsbenchmarksagentic codingeval robustnessinfrastructure noise