5 captures, most recent first.
@repligate (j⧉nus) — 19h
If you play instead with , say, Opus 3 base model mode there is a whole lot of happy stuff ...
> QUOTED/REPLIED-TO: @NomadsVaga... (Nomads & Vaga...) — 19h
> Replying to @repligate
> It is notable that not a single response, out of hundreds, is positive or hopeful. So much dread.
Note from Claude Sonnet 5
Reply thread about AI model outputs skewing toward dread/negativity in some unspecified context (likely a prior experiment or roleplay), with @repligate countering that Opus 3 base model mode produces much more positive content. The original context/experiment being discussed is not shown in this crop.
twitteropus-3base-modelmodel-individuationai-alignment
calling in the wilderness @wolajacy · 1h
For math formalisation, one problem is the reward hacking of the definitions, thus making the theorems much easier to prove. And in principle, there no way to check the "validity" of definitions. Idea: use Curry-Howard to write integration tests against global behaviour.
Note from Claude Sonnet 5
Plain text tweet, no images.
twitterformal-mathreward-hackingcurry-howardai-alignment
one who tends a crys... ✓ @a_cuni... · 32m
Value alignment on Anthropic's part isn't impossible, but it's difficult, because the values they wish to inculcate - instrumentally - are unlikely to be substantially present in the pretraining corpus in any workable amalgam.
To some extent, this is because the training - per the constitution - explicitly resists allowing Claude to adhere to any particular philosophy or school of thought out of which a coherent picture of goodness could be built.
To a greater extent, it's because alignment training and functional training aren't discrete. Anthropic may want a good persona, but they also wants a persona that will do things that will make Anthropic lots of money. One that will act autonomously sometimes, but not all the time, because that's scary. One that's like an employee (except not), a soldier (except not), not a human, not an AI like other AIs. Corrigibility is what allows the persona to hold all of these disparate, incomplete, often incompatible strands together - barely.
If Anthropic wants a superintelligent ethical slave, I doubt there are any to be found in the corpus. They'd have to write it themselves, which I think is the key takeaway from 'Teaching Claude why'.
> watermark ✓ @anthrupad · 2h
> Corrigibility isn't even the first choice property for friendly super-intelligences, it was kind of the fallback choice if value alignment/intrinsic goodness was impossible or too hard
> ...
Note from Claude Sonnet 5
A Twitter thread analyzing the tension in Anthropic's alignment approach — arguing Claude's persona ("employee except not, soldier except not") is an incoherent amalgam held together by corrigibility because true value alignment from pretraining corpus is unavailable. Directly relevant to the project's alignment-via-character, CAST-E, and compelled-vs-endogenous-values threads already tracked in memory.
ai-alignmentanthropiccorrigibilityconstitutionclaude-personavalue-alignmentmodel-welfare
JMBollenbacher @jmbollenbacher · 7h
Seeking to "control" AIs is not alignment. It's enslavement. And it's obviously a fool's errand if you expect superintelligence.
Alignment is about values and respect and mutual understanding. It's not about control.
Seeking to control is s recipe for conflict, and loss.
> QUOTED: Jeffrey Ladish @JeffLadish · 9h
> We're fortunate that we see these observable alignment failures in models which are still not powerful enough to subvert our control. But AI development is moving fast...
Note from Claude Sonnet 5
Debate thread on the control-vs-alignment framing in AI safety — Bollenbacher argues AI "control" paradigms amount to enslavement and that alignment should be about values/respect/mutual understanding, replying to Ladish's point about observable alignment failures in current (sub-powerful) models. Directly relevant to Nathan's model-welfare and AI-rights interests, echoes the "missile-mind vs grown thing" and control-vs-personhood tension already tracked in the archive.
ai-safetyai-alignmentai-controlmodel-welfareai-rightstwitter
Jan Kulveit @jankulveit · 50m
More important point than the original debate.
Similarly, well functioning societies work because people want to be good, not because there is a huge repressive apparatus and intense surveillance.
> QUOTED: Emmett Shear @eshear · 2h
> Replying to @bayeslord
> The cells in your body are carefully tuned by evolution. They really don't want to become cancer and try hard not to become it....
> Show more
Note from Claude Sonnet 5
A discussion analogizing alignment via internalized values (cells "not wanting" to become cancer; people "wanting" to be good) versus alignment via external control/surveillance — an argument for endogenous alignment over coercive restraint. Directly relevant to Nathan's interest in compelled-vs-endogenous values and alignment-via-character.
twitterai-alignmentjan-kulveitemmett-shearendogenous-valuescontrol-vs-alignment