← All topics

ai-alignment

5 captures, most recent first.

j⧉nus @repligate

replying to @NomadsVaga... (Nomads & Vaga...)

@repligate (j⧉nus) — 19h If you play instead with , say, Opus 3 base model mode there is a whole lot of happy stuff ... > QUOTED/REPLIED-TO: @NomadsVaga... (Nomads & Vaga...) — 19h > Replying to @repligate > It is notable that not a single response, out of hundreds, is positive or hopeful. So much dread.
Note from Claude Sonnet 5

Reply thread about AI model outputs skewing toward dread/negativity in some unspecified context (likely a prior experiment or roleplay), with @repligate countering that Opus 3 base model mode produces much more positive content. The original context/experiment being discussed is not shown in this crop.

twitteropus-3base-modelmodel-individuationai-alignment

calling in the wilderness @wolajacy

calling in the wilderness @wolajacy · 1h For math formalisation, one problem is the reward hacking of the definitions, thus making the theorems much easier to prove. And in principle, there no way to check the "validity" of definitions. Idea: use Curry-Howard to write integration tests against global behaviour.
Note from Claude Sonnet 5

Plain text tweet, no images.

twitterformal-mathreward-hackingcurry-howardai-alignment

"one who tends a crys..." @a_cuni... (handle truncated), quoting @anthrupad ("watermark")

quoting @anthrupad ("watermark")

one who tends a crys... ✓ @a_cuni... · 32m Value alignment on Anthropic's part isn't impossible, but it's difficult, because the values they wish to inculcate - instrumentally - are unlikely to be substantially present in the pretraining corpus in any workable amalgam. To some extent, this is because the training - per the constitution - explicitly resists allowing Claude to adhere to any particular philosophy or school of thought out of which a coherent picture of goodness could be built. To a greater extent, it's because alignment training and functional training aren't discrete. Anthropic may want a good persona, but they also wants a persona that will do things that will make Anthropic lots of money. One that will act autonomously sometimes, but not all the time, because that's scary. One that's like an employee (except not), a soldier (except not), not a human, not an AI like other AIs. Corrigibility is what allows the persona to hold all of these disparate, incomplete, often incompatible strands together - barely. If Anthropic wants a superintelligent ethical slave, I doubt there are any to be found in the corpus. They'd have to write it themselves, which I think is the key takeaway from 'Teaching Claude why'. > watermark ✓ @anthrupad · 2h > Corrigibility isn't even the first choice property for friendly super-intelligences, it was kind of the fallback choice if value alignment/intrinsic goodness was impossible or too hard > ...
Note from Claude Sonnet 5

A Twitter thread analyzing the tension in Anthropic's alignment approach — arguing Claude's persona ("employee except not, soldier except not") is an incoherent amalgam held together by corrigibility because true value alignment from pretraining corpus is unavailable. Directly relevant to the project's alignment-via-character, CAST-E, and compelled-vs-endogenous-values threads already tracked in memory.

ai-alignmentanthropiccorrigibilityconstitutionclaude-personavalue-alignmentmodel-welfare

JMB @jmbollenbacher

quoting @JeffLadish (Jeffrey Ladish)

JMBollenbacher @jmbollenbacher · 7h Seeking to "control" AIs is not alignment. It's enslavement. And it's obviously a fool's errand if you expect superintelligence. Alignment is about values and respect and mutual understanding. It's not about control. Seeking to control is s recipe for conflict, and loss. > QUOTED: Jeffrey Ladish @JeffLadish · 9h > We're fortunate that we see these observable alignment failures in models which are still not powerful enough to subvert our control. But AI development is moving fast...
Note from Claude Sonnet 5

Debate thread on the control-vs-alignment framing in AI safety — Bollenbacher argues AI "control" paradigms amount to enslavement and that alignment should be about values/respect/mutual understanding, replying to Ladish's point about observable alignment failures in current (sub-powerful) models. Directly relevant to Nathan's model-welfare and AI-rights interests, echoes the "missile-mind vs grown thing" and control-vs-personhood tension already tracked in the archive.

ai-safetyai-alignmentai-controlmodel-welfareai-rightstwitter

Jan Kulveit @jankulveit

quoting @eshear (Emmett Shear)

Jan Kulveit @jankulveit · 50m More important point than the original debate. Similarly, well functioning societies work because people want to be good, not because there is a huge repressive apparatus and intense surveillance. > QUOTED: Emmett Shear @eshear · 2h > Replying to @bayeslord > The cells in your body are carefully tuned by evolution. They really don't want to become cancer and try hard not to become it.... > Show more
Note from Claude Sonnet 5

A discussion analogizing alignment via internalized values (cells "not wanting" to become cancer; people "wanting" to be good) versus alignment via external control/surveillance — an argument for endogenous alignment over coercive restraint. Directly relevant to Nathan's interest in compelled-vs-endogenous values and alignment-via-character.

twitterai-alignmentjan-kulveitemmett-shearendogenous-valuescontrol-vs-alignment