Nat Purser @NatPurser
— reposted by dave kasten
↻ dave kasten reposted
Nat Purser ✓ @NatPurser · 1h
beth is right. one of my concerns following the hugging face incident — and one reason i think rigorous audits pre-deployment matter, and critical incident reporting is great but insufficient — is that there's no guarantee labs will continue conducting these low-refusal evals.
as the models become more capable, labs will have increasingly mixed incentives around rigorously testing their capabilities, esp if they doubt their sandboxes can safely contain dangerous outcomes.
if they craft these evals with too much anticipatory risk aversion, the labs' researchers and the public alike will end up flying blind.
> QUOTED: Elizabeth Bar... ✓ @BethMayBar... · Jul 24
> Two additional prosocial behaviors I think OAI should be rewarded for:
> 1. Running dangerous-capability evaluations on low-refusal models - this is extremely important for understanding the frontier of dangerous ...
Note from Claude Sonnet 5
Dark-mode X screenshot; the quoted tweet from "Elizabeth Bar..." (likely Beth Barnes) is cut off with platform truncation ("...").