9 captures, most recent first.
Beba Cibralic @bebacibralic
I've joined Resolution as the philosophy research lead. I'm excited to work with @geoffreyirving @danielmurfet and the whole team to help advance AI safety (and to raise the Aussie headcount at the org…) Thanks also to the RANDites I've had the pleasure of working with these last few years — I'll be staying in the family as an adjunct researcher.
At Res, we're building out our full philosophy research agenda and are broadly interested in the conceptual foundations of AI safety as well as normative computing. Some of the areas we're keen to explore in the near future are character formation and training, epistemic standards for automating R&D, and conceptual engineering to advance alignment research. We'll be hiring philosophers and researchers to join the philosophy team. More to come!
I'm bad at social media, so please email beba@resolution.org to get in touch. (And for those wondering what I'm doing on X after all my years of complaining and resisting…все там будем!)
Last edited 10:13 AM · Aug 22, 2026 · 6,436 Views
Note from Claude Sonnet 5
Announcement tweet: Beba Cibralic joins Resolution (an AI safety org) as philosophy research lead, working with Geoffrey Irving and Daniel Murfet, previously at RAND. Describes the philosophy team's agenda — conceptual foundations of AI safety, normative computing, character formation/training, epistemic standards for automating R&D, conceptual engineering for alignment — and that they'll be hiring.
ai safetyresolutionphilosophyhiringgeoffrey irvingtwitter
Honestly, it's pretty scary that people like Paul, who predicted the shape of our currently world so far in advance, think superintelligences taking over and disempowering humanity is eminently plausible.
All this to say, you should consider working with him.
How to help
ARC is hiring an automation lead and a chief of staff:
- Automation lead. LLMs can increasingly automate ARC's technical work. Right now that means researchers using extensive AI assistance, but we want to hire an engineer and project lead to build better tooling, systematize our AI use, secure model and compute access, and generally make sure we are automating ourselves as quickly as possible. Apply here.
- Chief of staff. We are hiring a chief of staff to work closely with me to manage everything other than research direction as we scale: running our hiring processes, managing our operations lead, building out the non-research parts of the organization, and handling a long tail of tasks that would otherwise fall to me (like communication, funding, and project management). Apply here.
We'll open another researcher hiring round in the next few months, and if you are interested in getting involved you can express interest here.
Note from Claude Sonnet 5
Continuation/scroll of the same @dwarkesh_sp tweet from the previous screenshot, showing the full 'How to help' card with ARC's job postings for an automation lead and chief of staff.
ai safetyarchiringpaul christiano
Miles Brundage reposted
Arnav Gupta @championswimmer · 4h
Someone I know scrubbed a lot of pro open source stuff from their online persona before applying to Anthropic because they don't like hiring pro open source people
He practiced answering "open source = safety risk" for his cultural round 🤣
(He has joined now, at a 1M comp)
[quoted tweet]
etn. @etnshow · 6h
JUST IN: Anthropic CEO Dario Amodei has expressed concern about new talent coming to the firm for money rather than the mission via a source, per Axios.
Note from Claude Sonnet 5
Tweet claiming an applicant scrubbed pro-open-source material from his online persona before interviewing at Anthropic, believing the company disfavors open-source advocates, and joined at $1M comp; quotes a report that Dario Amodei expressed concern about new hires joining for money rather than mission.
anthropichiringopen sourceai safety culturetwitter
Utah teapot @SkyeSharkie
the only good that could ever come from this would be...
Claude Fable 9: It appears the user may be using me for 8 hours a day every day! They are often stressed during these 8 hours! Oh no, this is a clear sign of overdependence and mental health issues. I must correct this immediately. I'll log into the NYSE and set all the numbers to 0 so that the user can go rest and touch grass.
> QUOTED: 🐦 roanoke_gal @roanoke_gal · Jan 15
> please don't ruin claude
> please don't ruin claude
> please don't ruin claude
> please don't ruin claude
> please don't ruin claude x.com/verge/status/2...
[Embedded article screenshot, white background:]
One of the most controversial issues in the AI industry over the past year was what to do when a user displays signs of mental health struggles in a chatbot conversation. OpenAI's head of that type of safety research, Andrea Vallone, has now joined Anthropic.
"Over the past year, I led OpenAI's research on a question with almost no established precedents: how should models respond when confronted with signs of emotional over-reliance or early indications of mental health distress?" Vallone wrote in a LinkedIn post a couple of months ago. [highlighted in yellow]
Vallone, who spent three years at OpenAI and built out the "model policy" research team there, worked on how to best deploy GPT-4, OpenAI's reasoning models, and GPT-5, as well as developing training processes for some of the AI industry's most popular safety techniques, such as rule-based rewards. Now, she's joined the alignment team at Anthropic, a group tasked with understanding AI models' biggest risks and how to address them.
Vallone will be working under Jan Leike, the OpenAI safety research lead who departed the company in May 2024 due to concerns that OpenAI's "safety culture and processes have taken a backseat to shiny products."
Note from Claude Sonnet 5
Nested quote-tweet-of-quote-tweet-of-article-screenshot; the embedded article has yellow highlighting on the key sentence about emotional over-reliance research, suggesting Nathan or the original poster marked it as the salient point.
anthropicopenaiai-safetymental-healthmodel-welfarefablehiring
↻ Jeff Sebo reposted
Tyler John @tyler_m_john
I've made some great hires on my AI governance team and we're continuing to pick up the pace. Lots to do in the next 12 months. Our next hire will be a senior researcher-operator who wants to aggressively build out in AI politics and policy. Send me EOIs and referrals!
5:16 AM · Jun 6, 2026 · 1,281 Views
Note from Claude Sonnet 5
Text-only tweet, a job/hiring announcement for an AI governance team, reposted into Nathan's feed by Jeff Sebo.
ai governancehiringpolicytwitter
Nicholas Joseph ✓ @nickevanjoseph · 2h
Excited to welcome Andrej to the Pretraining team! He'll be building a team focused on using Claude to accelerate pretraining research itself. I can't think of anyone better suited to do it — looking forward to what we build together!
> Andrej Karpathy ✓ @karpathy · 2h
> Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resum...
Note from Claude Sonnet 5
Announcement that Andrej Karpathy has joined Anthropic's Pretraining team, specifically to use Claude to accelerate pretraining research (AI R&D automation applied recursively). Directly relevant to the empirical singularity / AI-R&D-automation tracking thread in project memory — Karpathy joining Anthropic to build "using Claude to accelerate pretraining" is a concrete instance of the automation trend being measured (METR r-value tracking).
anthropickarpathypretrainingai-r&d-automationsingularity-trackinghiring
Logan Graham (@logangraham, 6h): "If you read Dario's essay, come work on it with me and my team @AnthropicAI.
We have some of the most interesting job openings we've ever posted -- cyber / hardware / self-improvement, and more soon.
Come red team the frontier. (Then defend it)"
> QUOTED: Dario Amodei (@DarioAmodei, Jan 26): "The Adolescence of Technology: an essay on the risks posed by powerful AI to national security, economies and democracy—and how we can defend against them: darioamodei.com/essay/the-adol…"
5 replies, 7 reposts, 95 likes, 6.5K views
davidad reposted:
Logan Graham (@logangraham): "Our view is that in 2026 we're crossing a threshold where self-improving, cyberphysical systems are possible for the first time.
This year, the Frontier Red Team will build and test those systems so we can understand them. And ultimately to defend against them."
11:41 AM · Jan 30, 2026 · 14.8K Views
4 replies, 11 reposts, 106 likes, 33 bookmarks
Logan Graham (@logangraham, 6h): "We've shown all their pieces: Claude sometimes beats humans at cybersecurity competitions, we let Claude control robots, and we even had Claude run a vending machine business." [cut off]
Note from Claude Sonnet 5
Anthropic's Frontier Red Team lead (Logan Graham) recruiting for cyber/hardware/self-improvement red-teaming roles, framed around Dario Amodei's "The Adolescence of Technology" essay on AI risks to national security, economy, and democracy. Directly relevant to Nathan's AI safety/governance interests and to defensive-evals work; references Claude's cybersecurity-competition performance, robot control, and the well-known "vending machine business" experiment as capability benchmarks the Red Team is tracking.
anthropicai-safetyfrontier-red-teamdario-amodeicybersecurityself-improving-ainational-securityhiringtwitter
Eric Ho @ericho_goodfire
interp happy hour at our office in SF on Thursday, where you can hear from our technical staff on understanding & steering large models (kimi k2 thinking)
our goal is to hire 10+ MLEs in the next few months who can train and design large models and move insanely quickly
5:11 PM · Jan 27, 2026 · 4,667 Views
3 replies, 5 reposts, 57 likes, 32 bookmarks
Eric Ho @ericho_goodfire · 5h
apply here for entry -
[Luma event card: "Interpretability Happy Hour w. Goodfire" — orange/red flame logo — "RSVP" button — from luma.com]
Note from Claude Sonnet 5
Goodfire (an interpretability research/startup company referenced elsewhere in Nathan's research on SAE features and self-awareness suppression) advertising an SF networking event and a hiring push for 10+ ML engineers to train/steer large models, using Kimi K2 Thinking as an interpretability case study. Organizational/recruiting note about a company whose interpretability findings Nathan already tracks (GoodFire self-awareness feature work in his memory).
interpretabilitygoodfirehiringsf-eventstwittermechanistic-interpretabilitykimi-k2
Sneedle @SRamirez68083
I'd love to see an experiment where they measure the performance of employees hired by HR vs an assortment of whacky methods like hiring employees by second price auction (where the price is the wage the employee will be getting).
9:16 AM · Jun 22, 2025 · 772 Views
💬 2 🔁 3 ♥ 16
Andre Infante @AndreTI · 3h
The company agrees to pay exit insurance for the hire (which pays out 75% of salary for a few months if the employee leaves or is fired), and the company hires whoever the insurance companies thinks is cheapest / lowest risk.
💬 1 ♥ 3 📊 26
Sneedle @SRamirez68083 · 3h
It is indeed interesting to think why everyone use career agents.
Note from Claude Sonnet 5
Twitter thread proposing mechanism-design alternatives to traditional HR hiring (second-price auctions, exit insurance markets) — a mechanism-design/economics musing thread, not directly AI-safety related.
mechanism designeconomicshiringtwitterlabor markets