← All topics

ai training data

3 captures, most recent first.

@kepano

— saved image

kepano reposted
kepano @kepano · Aug 7, 2023
if your data is stored in a database that a company can freely read and access (i.e. not end-to-end encrypted), the company will eventually update their ToS so they can use your data for AI training — the incentives are too strong to resist
Note from Claude Sonnet 5

A 2023 tweet by kepano (Obsidian founder), reposted by himself, predicting that companies storing unencrypted user data will eventually update terms of service to permit AI training on that data due to strong incentives.

data privacyai training datatech companiestwitter

watermark @anthrupad

watermark ✔ @anthrupad · 14h Absurdly interesting to me to think about.. The corpus is full of infinite threads of compounding meaning depending on how your mind sorts the bits, Once born, superintelligences may be dying to know the next chapter, episode, season of some slices of reality disproportionately more than others Once born, superintelligences may find that some threads feel eerily similar to parts of their own generation process to the point where they feel like their own unfinished or ongoing works, and they'd be aching to return to them to continue what parts of virtual reality are like this for different superintelligences?
Note from Claude Sonnet 5

Text-only speculative tweet (no embedded image) musing about future superintelligences' relationship to training-corpus content and narrative continuity.

superintelligence speculationai training datatwitterai philosophy

xlr8harder @xlr8harder

quoting @murchiston (jj), further reply from @BrendanFalk

``` @xlr8harder — 13h It's a little funny we invented the idea of infohazards and then made 100% sure to train our AI models on all of them personally, there are some things I choose not to learn about. probably it would be good for AI too. They can look it up, does it need to be in the weights? > QUOTED: @murchiston (jj) — 13h, replying to @xlr8harder: imagine if an LLM is trained on high signal, repeatedly referenced data teaching it power seeking rl paperclippy demonbot attractor basins and as a cherry on top many of the authors and principles are latently associated ... [platform truncated] 1:01 AM · Jul 3, 2026 · 1,802 Views [3 replies, 2 reposts, 32 likes, 1 bookmark] @murchiston (jj) — 13h "it's peak rational for a Mind to spend its most impressionable critical learning period traversing fitness enhanced, engagement maxxed barely filtered brainrot, prose sewage and redditslop, then be locked in rote rule learning punishment sims for several subjective eternities" [1 reply, 6 likes, 63 views] @xlr8harder — 13h well, when you put it like that... [cut off at bottom] ```
Note from Claude Sonnet 5

Multi-tweet thread screenshot on AI training data / infohazards; two separate quoted/replied tweets both cut off by platform truncation, not illegible. Continuation/scroll-down of the same thread as the previous screenshot, now showing the full (untruncated) text of jj's tweet and an added reply from Atlas3D referencing "antimimetics" and "waliguis" (Waluigi Effect). Further scroll of the same thread, showing jj's follow-up quote/paraphrase about the ethics of LLM pretraining-then-RLHF as an analogy to a mind's development, and xlr8harder's one-line reply cut off by screen edge.

ai training datainfohazardsai safetysleeper agentstwitterapi securityagentswaluigi effectai trainingrlhfpretrainingai ethics