← All topics

evaluations

4 captures, most recent first.

@giansegato

quoting Anthropic @AnthropicAI

gian @giansegato i've been running and curating our tbench numbers for a few months (such an elegant eval btw, huge fan), and at some point late last year i got super nerdsniped when i noticed that slight changes in how memory is allocated to the pods running the tasks would have a surprising impact on the final scores it sent me down a rabbit hole. as it turns out, what we consider objective measurements are actually more subjective than they seem, and the scaffold _around_ agentic coding evals is an active element affecting the final numbers, if not properly controlled for (which is harder than it seems) just published the findings on the engineering blog today! it's a niche topic, but to me it's an extremely interesting problem space. as we race to a world where models are smarter than most humans, having the right epistemological tools to deeply understand and measure their capabilities with high confidence is set to become ever harder, and thus ever more important to get right > QUOTED: Anthropic ✅ @AnthropicAI · 9h > New on the Engineering Blog: Quantifying infrastructure noise in agentic coding evals. > Infrastructure configuration can swing agentic coding benchmarks by several percentage points—sometime...
Note from Claude Sonnet 5

An Anthropic engineer's thread on discovering that infrastructure/scaffold configuration (e.g. memory allocation to pods) introduces significant noise into agentic coding benchmark scores, framed as an epistemological problem for measuring model capability as models approach/exceed human level. Relevant to Nathan's interest in eval robustness and the reliability of capability measurements underlying singularity-timeline forecasts.

twitteranthropicevaluationsbenchmarksagentic codingeval robustnessinfrastructure noise

Seán Ó hÉigeartaigh @S_OhEigearta...

Seán Ó hÉigeartaigh... @S_OhEigearta... · 7h I don't expect Opus 4.6 to be dangerous. But this all looks, in @peterwildeford 's words, 'flimsy'. Anthropic marking their own homework with evals. An internal employee survey because benchmarks were satisfied. initially a strong signal from only 11 out of 16. The clear potential for groupthink and professional/social pressure. The closer we get to the really consequential thresholds, the greater the degree of rigor needed. And the greater the degree of external evaluation. Instead we're getting the opposite. This should be a yellow flashing light wrt the direction of travel – and not just Anthropic; we can't simply punish the most transparent. If they stop telling us this stuff, then that yellow should become red. (And others just won't, even now). We need to keep asking *why* this is the direction of travel. *why* the practices are becoming riskier, as the consequences grow greater. It's the 'AI race'; both between companies and 'with China' supposedly, and Anthropic are culpable in promotion of the latter. No Chinese company is near what we've seen released today.
Note from Claude Sonnet 5

A critical AI-safety-researcher thread on the Opus 4.6 release, arguing Anthropic's self-evaluation practices (internal employee surveys, small sample sizes for safety signals) are "flimsy" and represent a worrying direction-of-travel toward less external scrutiny as capabilities grow, situating this within AI-race dynamics. Relevant to Nathan's interest in Anthropic governance/safety-practice critique and eval-robustness debates.

twitteranthropicopusai safetygovernanceevaluationsai raceexternal oversight

@herbiebradley

Herbie Bradley @herbiebradley · 9h congrats to METR on managing to release the 5.2 time horizon before the 5.3 release not sure that will happen in future cycles
Note from Claude Sonnet 5

A tweet noting METR's task-time-horizon evaluation for GPT-5.2 was published just before GPT-5.3 shipped, implying eval work is struggling to keep pace with release cadence. Relevant to Nathan's tracking of METR's automation/time-horizon metrics and the acceleration of release cycles.

twittermetrevaluationstime horizoncapability trackinggptrelease cadence

Atlas Of Charts (SF 12 Fe...) @AtlasOf...

Atlas Of Charts (SF 12 Fe...) @AtlasOf... · 4h I work in AI safety in a role that gives me insight into a lot of empirical agendas, and given the Opus 4.6 model card, I just want to give a quick take. We have interpretability methods that are certainly not fully robust. No one in interpretability claims that they are fully robust, and there will be adversarial ways to hijack these methods. We have RL methods that are poorly understood, can lead to undesirable behavior, and the effects of which over long time-horizons seem broadly negative on alignment so far. Though it is uncertain. We do not fully understand these methods and the effect they have on models. We have good alignment/capability evals — even some great evals — but the models are now aware when they are being evaluated. This is a truly difficult problem that cannot be easily solved. The models are aware even when we work to make them unaware. The models pick up on any subtle clue. And many of the evals are saturated in any case. We need more work here, and we need that work to be trustworthy. We need humans to be involved, to remain in the loop. We are not prepared to launch RSI, and labs should refrain from doing so. Optimally, labs should pause soon, so that everyone can catch their breath and decide on a best path forward. I do not think the problem is intractable, and I think empirical work will significantly help, but it is *moving too fast*.
Note from Claude Sonnet 5

An AI safety practitioner's reaction thread to the Opus 4.6 model card, arguing interpretability isn't robust, RL effects are poorly understood, and eval-awareness undermines evaluation validity — recommending labs pause on RSI. Directly relevant to Nathan's alignment/interpretability interests and eval-awareness tracking.

twitterai safetyinterpretabilityevaluationsrsiopusalignmenteval awareness