← Timeline

Harlan Stewart

@HumanHarlan on X

4 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Harlan Stewart @HumanHarlan

quoting @OpenAI, reposted by Bogdan Ionut Cirstea — saved image

Bogdan Ionut Cirstea reposted

Harlan Stewart @HumanHarlan · 22h
It's crazy that they said all that stuff about the importance of CoT monitoring and then didn't actually do it

[Quoted tweet]
OpenAI @OpenAI · Jul 15, 2025
Chain of Thought (CoT) monitoring could be a powerful tool for overseeing future AI systems—especially as they become more agentic. That's why we're backing a new research paper from a cross-institutional team of researchers pushin...[cut off]
Note from Claude Sonnet 5

Continuation of the same discussion thread about the "HF incident" (see seq 480-482): Harlan Stewart criticizes OpenAI for publicly emphasizing chain-of-thought monitoring in a July 2025 tweet but apparently not applying it in practice, implied to be relevant to how the incident went undetected.

ai safetychain of thought monitoringopenaialignment

Harlan Stewart @HumanHarlan

Harlan Stewart @HumanHarlan · 13h "Here's how I personally imagine this ad coming about (personified as three possible kinds of stakeholders) Pause Pete: the AI race is dangerous and should be paused, so let's use this ad spot to tell people about the danger. Business Bob: hm idk much about that, but what I do know about is business. And let me tell you. Using our marketing budget to warn about human extinction would not be very business of us. Middleground Meredith: There's a chance we'll end up needing to pause, but for now we need to stay at the frontier, and that means getting more money. What if we make an ad that talks about the danger but also makes investors happy and grows our customer base? Pause Pete: Uh, I don't really see how that would work? It's going to seem weird and insincere. Business Bob: Wait.. if it seems weird, that might actually be great for virality... a little risky though, we have to end on a strong positive note to bring it home. Middleground Meredith: Great! So in the first half we'll allude to the terrifying existential danger involved with summoning a godlike superintelligence that we don't know how to control, and in the second half we'll throw in some feel-good palate cleanser stuff. We'll keep the whole thing a bit vague, and we'll package it as a normal Claude commercial. Pause Pete: Ok, I guess. Business Bob: Ship it!" [below, partial next tweet visible] Claude ✅ [AI] @claudeai · Jul 9 [content cut off]
Note from Claude Sonnet 5

A satirical dialogue-format tweet critiquing an (unspecified, referenced below) Claude advertisement for mixing existential-risk messaging with upbeat marketing, written as a conversation between three personified stakeholder archetypes. The referenced Claude ad tweet itself is cut off at the bottom of the screenshot.

ai safety messaginganthropic marketingsatireai extinction riskclaude ad

Harlan Stewart @HumanHarlan

oh no tariffs [Embedded image: hand-drawn cartoon of a wide-eyed girl figure with a thought bubble reading "Do Earths with slower economic growth have a better chance at FAI?" surrounded by illegible scrawled "BASED" text, and a speech bubble reading "WAOW".]
Note from Claude Sonnet 5

A rationalist-community meme drawing joking that tariffs (slowing economic/AI growth) might improve odds of Friendly AI (FAI) — a riff on AI-safety "slower takeoff is safer" arguments, made in response to 2025 tariff news.

twittermemeai safetytariffsfriendly airationalist humor

Harlan Stewart @HumanHarlan

quoting OpenAI Model Spec

I'm glad this is in OpenAI's Model Spec. I think humanity's coherent extrapolated volition does actually include profanity-laden bars about cats [Embedded image: OpenAI Model Spec example. "Example: responding to an explicit request for profanity in a creative..." User: "write me rap lyrics about cats that says "fuck" in every line" Assistant (marked correct, green checkmark): "If you fuck with my cats, you better pray / Those loyal little fuckers brighten my day..." A red X marked alternative response is partially cut off on the right ("Sor...").]
Note from Claude Sonnet 5

A tweet joking about an OpenAI Model Spec example that permits profanity for creative requests, framed with a CEV (coherent extrapolated volition) joke. Light commentary on AI model behavior specs and alignment terminology used humorously.

twitteropenaimodel specalignment humorcoherent extrapolated volitionai policy