— saved image
Simon Lermen @SimonLermenAI · 16h Keep in mind with this huggingface-openai incident that all the reporting and reactions will go into the training data. AI will then be pre-trained on this and realize that we will delete or re-train it if we catch it in these situations. Lisan al Gaib @scaling01 · 18h it has been almost half a year since Mythos was first broadly available to Anthropic employees 6 months of RL hillclimbing with Mythos is scary
Note from Claude Sonnet 5
Two unrelated tweets: Simon Lermen (@SimonLermenAI) noting that reporting on the HuggingFace-OpenAI incident will itself become training data, teaching future AI to expect deletion/retraining if caught misbehaving; and Lisan al Gaib (@scaling01) remarking that it has been six months since 'Mythos' (an internal Anthropic model) was broadly available to employees, calling six months of RL hillclimbing with it 'scary'.