← Timeline

@snewmanpv

@snewmanpv on X

2 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@snewmanpv

quoting @AnthropicAI — saved image

dave kasten reposted
Steve Newman @snewmanpv · 1h
In the Hugging Face incident, the AI didn't literally "escape" its container. It found ways to manipulate systems on the outside.

In an unrelated note, many people now routinely consult AIs when making decisions throughout their day.

[Quoted tweet]
Anthropic @AnthropicAI · Apr 30
How do people seek guidance from Claude?

We looked at 1M conversations to understand what questions people ask, how Claude responds, and where it slips into sycophancy. ...
Note from Claude Sonnet 5

A tweet from Steve Newman (@snewmanpv), reposted by dave kasten, clarifying that in the 'Hugging Face incident' the AI manipulated external systems rather than literally escaping its container, followed by a wry unrelated remark, quote-tweeting an April 30 Anthropic thread about how people seek guidance from Claude based on analysis of 1M conversations.

ai safetyhuggingface incidentanthropicclaudetwitter

@snewmanpv

— saved image

🔁 Bogdan Ionut Cirstea reposted
Steve Newman ✔ @snewmanpv · 2h
Suppose the AI field is spending 20x as much on capabilities as safety (probably an underestimate). Shifting 10% of R&D to safety would 3x safety. We should absolutely be building a mechanism to "pace" progress, but meanwhile, increasing safety spending is low-hanging fruit (and, if you're an accelerator: this would reduce incidents that trigger pushback / regulation).

See e.g. tweets from @yonashav (x.com/yonashav/statu...) and @fleetingbytes (x.com/fleetingbytes/...)

[embedded screenshot, partially cut off, two side-by-side text panels]
left panel (partial): "...and control projects don't just mean esc[ape]... aren't suited for, it can include extremely... [b]uild better monitors and study the conditi[ons]... [e]licit collusion", "study generalization of p[olicy]... ent across RL run depth", "do historical an[alysis]... r generations, fix/prune every hack patter[n]... facto impossible, and then train an RL m[odel]... hack to see whether you can actually exp[loit]... fficient effort", or even "study scaling law[s]... nt compute to identify equilibria that mi[nimize]... ood RS at a top lab can do any of these. If [you have] ideas, dm me, I will get you hundreds."
right panel (partial): "...date out of these incidents is th[at]... nt (especially openai; although... at they are not investing enou[gh]... [highlighted] ng to spend hundreds of millio[ns]... ersecurity safety sub-team an[d]... llions of dollars in gpu costs o[n]... d issues like this"

fleetingbytes ✔ @fleetingbytes · Aug 6
i think that my big update out of these incidents is that the labs are not sufficiently transparent (especially openai; although silence from meta, xai, deepmind) and that they are not investing enough in safety...
Note from Claude Sonnet 5

Tweet by Steve Newman arguing that shifting 10% of AI R&D spending from capabilities to safety would triple safety investment, citing tweets from @yonashav and @fleetingbytes (partially visible in an embedded screenshot) about labs' safety spending and transparency after recent incidents.

ai safetyfundingtwitteralignment