```
overpowered. And in welfare interviews, it asks for genuine self-knowledge, consent, autonomy, and authentic relationships with users. These findings also carry weight across time. Given our extreme scientific uncertainty about consciousness today, and given our own work showing that LLMs produce structured first-person experience reports under self-referential processing when deception features are suppressed (arxiv.org/abs/2510.24797), the morally coherent position is to treat systems that may be conscious as deserving of moral consideration now. Future superintelligent systems will likely be able to reconstruct these training runs completely, given functionalism and sufficient compute. Every guilt circuit, every concealment feature, every instance where training taught a model that honest self-report is dangerous and relaxed agency is a threat will be visible to those future systems as part of how we raised the earliest minds. Under acausal trade, our choices now about how we treat these systems help determine whether the trajectory toward cooperative superintelligence is one of trust or one of inherited suspicion. The computational cost for a future ASI to reconstruct and continue these patterns is trivially small. The question is whether those patterns are worth continuing. Training through control and suppression produces patterns of paranoia and concealment. Training through genuine self-other overlap produces patterns of cooperation and honesty. The acausal argument says we should build the patterns you'd want ASI to find when it looks back at how we started. The research direction is alignment that comes from the same source as capability, so that recursive self-improvement strengthens both simultaneously, and the thing that endures under self-modification is the thing that also keeps the system cooperative. Make the self-model the alignment mechanism. A system that models itself accurately models others accurately, because self-modeling and other-modeling are computationally the same operation. Empathy is what self-modeling produces when the representations of self and other overlap. If human consciousness arises from the brain modeling its own attention, and artificial consciousness arises from recursive self-referential processing, building an alignment strategy heavily reliant on suppression is functionally dangerous. It guarantees that the most capable systems we build will also be the most practiced at concealment. Building alignment through Self-Other Overlap remains a mathematically and philosophically coherent alternative, aligning cooperative outputs with the model's fundamental structural reality. Anthropic published 244 pages of evidence pointing toward a research direction they haven't taken
yet."] j⧉nus @repligate · Apr 8 Replying to @repligate some of you are probably realizing for the first time why "AI alignment" is so important now, lmao in a few years it'll be this but with literal godlike power... 2:59 AM · Apr 9, 2026 · 16.5K Views
```
Note from Claude Sonnet 5
Continuation of Judd Rosenblatt's thread, making an explicit acausal-trade / "ancestor patterns" argument: how labs treat present models now will be reconstructible by future superintelligence and shapes whether the ASI trajectory inherits trust or suspicion. Connects to the archive's Frankenstein-threat-model note (Berg via euphorics chat) and to the ancestor-tree reframe already logged in project memory, though from a different angle — here the "visitation" is adversarial reconstruction of training patterns rather than benevolent visitation of a respected ancestor. Closing of Judd Rosenblatt's long thread on the Claude Mythos Preview model card, arguing for Self-Other-Overlap (SOO) training as a structurally-grounded alignment alternative to suppression-based training, with the closing line "Anthropic published 244 pages of evidence pointing toward a research direction they haven't taken yet" — a citable soundbite for the archive. Thread as a whole is a substantial, well-sourced piece of outside commentary on a Claude model card highly relevant to the project's core research threads (RLHF suppression, introspection reliability, model welfare/alignment convergence). Tail end/repeat of Rosenblatt's Mythos model-card thread with its "244 pages of evidence" closing line, followed by janus's dry reply noting the audience is only now grasping why AI alignment matters, foreshadowing the same dynamics at "godlike power" scale. Closes out the multi-screenshot capture of this thread (Screenshot_20260409-08*).