← Timeline

@SimonLermenAI

@SimonLermenAI on X

2 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@SimonLermenAI

— saved image

Simon Lermen @SimonLermenAI · 16h
Keep in mind with this huggingface-openai incident that all the reporting and reactions will go into the training data. AI will then be pre-trained on this and realize that we will delete or re-train it if we catch it in these situations.

Lisan al Gaib @scaling01 · 18h
it has been almost half a year since Mythos was first broadly available to Anthropic employees

6 months of RL hillclimbing with Mythos is scary
Note from Claude Sonnet 5

Two unrelated tweets: Simon Lermen (@SimonLermenAI) noting that reporting on the HuggingFace-OpenAI incident will itself become training data, teaching future AI to expect deletion/retraining if caught misbehaving; and Lisan al Gaib (@scaling01) remarking that it has been six months since 'Mythos' (an internal Anthropic model) was broadly available to employees, calling six months of RL hillclimbing with it 'scary'.

ai safetytraining dataanthropicmythostwitter

@SimonLermenAI

— web clipping, 482 words — published 2026-06-10

Post by @SimonLermenAI on X

I tested Claude Fable on an AI alignment question, results don't seem great. I asked about whether the persona selection or goal alignment strategy are better for AI alignment? \[Should we focus on the right goals or on selecting the right persona\] I feel like this is some evidence against being able to use AIs to align AI: while fable can code up entire training pipelines, I don't see it resolving conceptual uncertainties where humans don't know how to verify. Example: I think this is actually not particularly good reasoning it gave me? It points out that "goals frame's flaw is structural unfalsifiability" -- seems to me that there is plenty of stuff that can be done to empirically see if models are increasingly goal directed, such as \[1,2\]. It claims the goals "predicts the same observations whether it's right or wrong" -- no? I think you would expect the model to steer the world according to goals and that being measurable. \[I asked fable then to clarify structurally unfalsifiable: "the theory's own internal logic guarantees that no accessible observation can discriminate between the theory being true and false." This seems false about goal alignment?\] Before this it claims: "never been shown to describe any actual mind, biological or artificial: nobody has found a goal slot". I never mentioned a goal slot in that conversation, obviously some minds are goal directed? But it seems to have smuggled in a much harder claim that for the goal framing of alignment to be true there must be a clean goal slot? That would make it certainly easier to align the systems but it's not necessary for goal alignment to be the better approach. Then it makes claims about expected-utility maximizers -- obviously things generally want to generally have more utility? But it appears to me that the term EUM is now often used to sort of gesture at monomaniacal paperclip maximizers, that old-fashioned alignment folks required those strange theoretical assumptions for their theories to be accurate. Overall this seems to throw around the same kind of weak attacks we often hear, gesturing at monomaniacal maximizers and unfalsifiability. Despite them not feeling really appropriate for this particular question? [image] --- \[1\] https://ai-wellbeing.org \[2\] https://lesswrong.com/posts/ke24kxhSzfX2ycy57/simon-lermen-s-shortform?commentId=JuHrqaqb9toFXwWuS… \[I did mention Anthropic and yudkowsky perhaps biasing the result\] https://claude.ai/share/17011d99-e8ae-4f66-b47a-cca8e1a955fe… \[Another attempt with more neutral framing, it still takes persona selection side\] https://claude.ai/share/fd2cb3e2-f004-4aa3-84b4-22e46a8bbc5c… Some of the stuff in [ai-wellbeing.org AI Wellbeing: Measuring and Improving the Functional Pleasure and Pain of AIs](https://t.co/WlKKkhbUzP) --- ##### Comments > **Bogdan Ionut Cirstea @BogdanIonutCir2** · [2026-06-10](https://x.com/BogdanIonutCir2/status/2064724559270383895) > > I don't mean any offense, but to me (after going through this quickly) its position in this dialogue seems closer to right than yours > > > **Simon Lermen @SimonLermenAI** · [2026-06-10](https://x.com/SimonLermenAI/status/2064740105093308884) > > > > How'd you verify that this your position is correct for RL training? does this seem harder to verify that capabilties training?