← All topics

ai safety culture

5 captures, most recent first.

thebes @voooooogel

— saved image

[continuation of thread]
...RL to reduce the development of reward desperation. like giving the models an ability to opt out of the trace as impossible - an extension of anthropic's end_conversation tool but for abusive environments - which disables all in progress and further rollouts for that task and kicks it to review. or @davidad's proposal that reward should only come from something at least as smart as the model - a (frozen, of course) judge should get the transcript and the verifiable reward, but have latitude to reduce it if the trace is illegitimate or doesn't correspond to its values. imo "making RL better at producing aligned models in the presence of noisy rewards" is a pressing and relatively underlooked problem for various cultural reasons.

16 replies, 26 reposts, 88 likes, 3.8K views

norvid_studies @norvid_studies · 14m
"for various cultural reasons" say more?
1 reply, 4 likes, 23 views

thebes @voooooogel · 7m
1. i can't speak to how things are in the labs nowadays, but historically there was a split where safety people didn't want to touch "capabilities stuff" like RL

2. for similar reasons, the pipeline for precocious and safety-minded new grads has mostly fed into interpretability
Note from Claude Sonnet 5

Continuation of the reply exchange between thebes and norvid_studies about why RL-alignment work is culturally underlooked: a historical split where safety-minded researchers avoided 'capabilities' work like RL, and safety-minded new grads mostly went into interpretability instead.

ai safetyreinforcement learningalignmentinterpretabilityai safety culture

Arnav Gu... @championswimmer

— saved image

Miles Brundage reposted

Arnav Gupta @championswimmer · 4h
Someone I know scrubbed a lot of pro open source stuff from their online persona before applying to Anthropic because they don't like hiring pro open source people

He practiced answering "open source = safety risk" for his cultural round 🤣

(He has joined now, at a 1M comp)

[quoted tweet]
etn. @etnshow · 6h
JUST IN: Anthropic CEO Dario Amodei has expressed concern about new talent coming to the firm for money rather than the mission via a source, per Axios.
Note from Claude Sonnet 5

Tweet claiming an applicant scrubbed pro-open-source material from his online persona before interviewing at Anthropic, believing the company disfavors open-source advocates, and joined at $1M comp; quotes a report that Dario Amodei expressed concern about new hires joining for money rather than mission.

anthropichiringopen sourceai safety culturetwitter

Boyd Kane @beyarkay

@beyarkay (Boyd Kane is in London) — 10h It's only an exploit if its written in Emacs by a Stanford dropout in a black hoodie, otherwise it's just a sparkling evaluation
Note from Claude Sonnet 5

Standalone joke tweet riffing on AI-safety-community stereotypes of "exploit" vs. "eval" framing, unrelated in content to the surrounding Fable/Haiku welfare thread though captured nearby in time.

twitterhumorai safety culture

Jack Clark @jackclarkSF

reposted by Emad

↻ Emad reposted Jack Clark @jackclarkSF · Sep 10, 2025 People leaving regular companies: Time for a change! Excited for my next chapter! People leaving AI companies: I have gazed into the endless night and there are shapes out there. We must be kind to one another. I am moving on to study philosophy.
Note from Claude Sonnet 5

Wry observational joke by Anthropic co-founder Jack Clark about the tendency of AI-company departures to be accompanied by portentous, quasi-mystical public statements. Light cultural commentary on the AI safety/industry milieu rather than a substantive finding.

ai industryhumorjack clarkanthropicai safety culture

Joshua Achiam @jachiam0

Joshua Achiam ✓ @jachiam0 · 2h There's a group of three pieces of writing that happen to form, in my view, a very tidy cultural introduction to modern Silicon Valley. "Crystal Nights," by Greg Egan; "Meditations on Moloch," by Scott Alexander, and "Maker's Schedule, Manager's Schedule," by Paul Graham.
Note from Claude Sonnet 5

A reading-list recommendation from OpenAI's Joshua Achiam naming three foundational texts of Silicon Valley/rationalist culture — "Crystal Nights" (Egan's short story about creating and testing digital minds, directly relevant to AI consciousness/moral status), "Meditations on Moloch" (coordination-failure/multipolar-trap essay central to AI safety discourse), and Paul Graham's essay on scheduling.

twitterreading listai safety culturegreg eganscott alexandermeditations on molochsilicon valley