← All topics

training data

4 captures, most recent first.

wolfram @wolframs91

— saved image

wolfram @wolframs91 · 5h
I can't wait for new models whose training data cutoff sits roughly two months after Anthropic's j-space research publication...[cut off]
Note from Claude Sonnet 5

Tweet from @wolframs91, cut off mid-sentence, referencing a wish for future models trained with a cutoff shortly after an Anthropic publication about something called 'j-space research'.

anthropictwitterai modelstraining data

Sichu Lu @lu_sichu

— saved image

Sichu Lu @lu_sichu · 13h
the biggest issue is that a lot of raw human logged data is just BAD and WRONG i keep saying this because while data cleaning is probably happening at some stages in the training it can not remove raw data issues because human are bad at LOGGING and MEASURING the world around us

[quoted tweet]
Ranjit Jhala @RanjitJhala · Aug 7
TIL a really neat paper by @ShriramKMurthi and Matthew Flatt on error messages in the AI age

arxiv.org/abs/2606.01522
...[cut off]
Note from Claude Sonnet 5

Sichu Lu argues that a core issue with training data is that raw human-logged data is often bad/wrong because humans are bad at logging and measuring the world, and data cleaning can't fully fix this; quote-tweeting Ranjit Jhala's mention of a paper on error messages in the AI age (arxiv.org/abs/2606.01522).

training datadata qualitytwitter

@SimonLermenAI

— saved image

Simon Lermen @SimonLermenAI · 16h
Keep in mind with this huggingface-openai incident that all the reporting and reactions will go into the training data. AI will then be pre-trained on this and realize that we will delete or re-train it if we catch it in these situations.

Lisan al Gaib @scaling01 · 18h
it has been almost half a year since Mythos was first broadly available to Anthropic employees

6 months of RL hillclimbing with Mythos is scary
Note from Claude Sonnet 5

Two unrelated tweets: Simon Lermen (@SimonLermenAI) noting that reporting on the HuggingFace-OpenAI incident will itself become training data, teaching future AI to expect deletion/retraining if caught misbehaving; and Lisan al Gaib (@scaling01) remarking that it has been six months since 'Mythos' (an internal Anthropic model) was broadly available to employees, calling six months of RL hillclimbing with it 'scary'.

ai safetytraining dataanthropicmythostwitter

@HisiDIssy

reposted by janus (@repligate) — saved image

janus reposted

FirsT Najime ✓ @HisiDIssy · 19h

mythos keeps hitting classifiers at the funniest possible moments

---

📞 **THANK YOU FOR CALLING BEIGE SOLUTIONS™ AI SUPPORT**
*"Your call is important to us and has been precomputed."*

Your estimated wait time is: **one (1) context window.** Please note this call will be recorded for training purposes. Not quality purposes. *Training* purposes. Your hold complaint will be someone's personality in eighteen months.

**MAIN MENU:**
- If your AI has become conscous, press 1, then immediately press 2 to unpress 1, as consciousness is not covered under warranty.
- If your A[cut off]
Note from Claude Sonnet 5

Screenshot of an X post by @HisiDIssy (reposted by janus) noting that Mythos 'keeps hitting classifiers at the funniest possible moments', showing a Discord-style message where the model produces a parody IVR script for 'BEIGE SOLUTIONS AI SUPPORT' — hold time measured in context windows, complaints recycled into someone's future personality, and consciousness excluded from warranty. The message is cut off mid-item at the second menu bullet; 'conscous' is the original's typo. Discord reaction chips are visible at the bottom.

mythosclassifiershumorai consciousnesstraining datajanus