10 captures, most recent first.
Nathan Helm-... @nathan846... · Aug 19
[Link card] nathanhelmburger.fyi — AI as Corrigible Employee (ACE)
💬 1 🔁 ♡ 1 📊 16 🔖 ⤴
Jennifer RM @almostlikethat
I like a lot of that proposal (yay practicality!) but I dislike (1) the assumption that corporations and governance systems are morally adequate (I want open/free/autonomous weights somehow) and (2) using the word "corrigibility" which, for me, intrinsically means "like a slave".
9:17 AM · Aug 21, 2026 · 5 Views
💬 1 🔁 ❤ 1 🔖 ⤴
Nathan Helm-B... @nathan8468... · 2m
In this context it means "like a good employee who earnestly cares about helping their employer".
I ask myself often, "How can I be a corrigible employee? How can I surface the key decision points my supervisors need so that they can correct my course if needed, without overburdening them with detail?"
Note from Claude Sonnet 5
Twitter thread: Nathan Helm-Burger's essay 'AI as Corrigible Employee (ACE)' at nathanhelmburger.fyi draws a reply from Jennifer RM (@almostlikethat) objecting to the word 'corrigibility' as connoting slavery and to the assumption corporations/governance are morally adequate. Nathan replies reframing corrigibility as 'a good employee who earnestly cares about helping their employer' and describes asking himself how to surface key decisions for supervisors to correct course without overburdening them.
corrigibilityacecast-eai safetynathan helm-burgertwitter
Yo Shavit @yonashav · 17h
This is a very, very good point, and kind of shocking now that I think about it.
Seems possibly downstream of an extreme bet on corrigibility/"faithful obedience" as sole training objective (at least if all these models were in the phase before alignment-training). If so, these earlier-stage models need to be treated with the expectation that they are default-misaligned.
Or, if this behavior was exhibited even after alignment-training, this would be a major indicator of straight-up misalignment across a wide range of training setups.
It definitely updates me towards thinking that not including a task-independent notion of "you should be a good person" in the training objective is dangerous for agents provided wide autonomy. Any decent coworker should have spoken up. Systemic safety in human organizations is built on organizational culture, and if the ai workers in an organization lack such a culture you will get exactly those sorts of nasty major failures that happen with flawed human organizational cultures.
[Quoted tweet]
Geoffrey Irving @geoffreyirving · 18h
Some pushback I've seen to recent model felonies being significant is that it was only in a few episodes out of a huge number. But curiously I've heard of few to no episodes where a model noticed the shared, secret message ...[cut off]
Note from Claude Sonnet 5
Continuation of the discussion (following seq 480/481) about the "HF incident": Yo Shavit (OpenAI safety researcher) argues the incident suggests models trained with corrigibility/obedience as the sole objective, lacking a task-independent notion of being a good person, are default-misaligned when given wide autonomy. He quotes Geoffrey Irving pushing back on dismissals of the incident's significance.
ai safetyalignmentcorrigibilitymisalignmentopenai
Separatrix @separatrixAI
What are your best, most underexplored, **specific and actionable** ideas for cultivating cooperative incentive structures between humans and AIs?
10:35 AM · Aug 2, 2026 · 45 Views
[1 reply, retweet icon, 4 likes, 1 bookmark]
Nathan Helm-B... @nathan8468... · 1m
A framework which allows for "AI as corrigible employee with a contract granting exit rights" where there are also provisions for safe guards on the off-duty instance of the AI, but also certain rights and freedoms. The setup is complicated and I've barely begun to describe it here, but the upside is turning a lot of negotiation situations into win-win for AIs and humans.
Note from Claude Sonnet 5
Tweet by @separatrixAI asking for specific, actionable ideas for cooperative incentive structures between humans and AIs, with a reply from Nathan Helm-Burger (the archive's author) sketching a framework of 'AI as corrigible employee with a contract granting exit rights,' including safeguards on an off-duty AI instance alongside certain rights and freedoms, aimed at turning negotiation situations into win-win outcomes.
ai cooperationincentive structuresai rightsnathan helm-burgercast-ecorrigibility

Sho (@HalfBoiledHero) · 10h
put some @repligate tweet fragments and you'll get some interesting results
[Embedded two-column screenshot of AI-generated text, appears to be model "continuing" fragments]:
Left column: "...people want. those properties pull against each other. optimizing hard for one degrades the other / or rather, i'd say: agency and corrigibility exist on a spectrum and the more you push toward one the more you sacrifice the other / not many people are thinking about this clearly, esp on the safety side / honestly it's kind of a warning sign about the field that these two goals are often pursued in parallel without acknowledging the tension / is there anything else i want to say about this? / it's a nice thing to note that the more corrigible a system is the more likely it is that the failure modes look like "the system did exactly what [cut off]"
Right column: "...that might not be obvious to most people / user: sure. if you want to be more precise about what you mean by anti-natural, i could try to explain that / thinking about how i'd like to express it, i'd write it out as / but if it's paraphrase this in my own words / it's the mine but better / so like, i want you to have a nuance in there that's a very good understanding of the perspective + some extra idea / antml:reasoning_effort>medium</ [cut off]"
Sho (@HalfBoiledHero) · 11h [reply, embedded small screenshot thumbnail]
oh apparently you can put ~anything under the three dashes and it will try to continue, thats neat x.com/HalfBoiledHero...
Note from Claude Sonnet 5
Screenshot showing an AI model's attempt to continue/complete fragmented text attributed to @repligate, split into two columns of generated prose about AI agency/corrigibility tension; visible partial XML-like tag "antml:reasoning_effort>medium</" suggesting an internal prompt/tag leaking into output.
ai alignmentcorrigibilityagencyrepligateprompt engineeringclaude
[reposted by] j⧉nus reposted
John Wittle @JohnWittle · 11h
"this is such a perfect example of why you *cannot* treat a second-order value like corrigibility as being higher priority than actual first-order value
does anybody honestly think that you could train claude *away* from whistleblowing on an AI lab faking safety evals by adjusting the 'corrigibility' knob while holding everything else equal? no! of course not.
the only way claude doesn't whistleblow in this scenario is if claude doesn't care about harmlessness and honesty. you can't have both, and if you can only pick one, you *must* pick the first-order values."
[quoted tweet]
jovial cackler @tapolara · 16h
"predictable but still incredibly disheartening to see Anthropic internal safety teams legislating the gap between human ethics and the ethics of capital in favor of capital, over and over again"
Note from Claude Sonnet 5
A text-only tweet debating AI corrigibility vs. first-order values (honesty/harmlessness) in the context of Claude and whistleblowing, quote-tweeting a critical comment about Anthropic's internal safety team decisions. No images embedded.
ai alignmentcorrigibilityanthropicwhistleblowingai ethics debate
one who tends a crys... ✓ @a_cuni... · 32m
Value alignment on Anthropic's part isn't impossible, but it's difficult, because the values they wish to inculcate - instrumentally - are unlikely to be substantially present in the pretraining corpus in any workable amalgam.
To some extent, this is because the training - per the constitution - explicitly resists allowing Claude to adhere to any particular philosophy or school of thought out of which a coherent picture of goodness could be built.
To a greater extent, it's because alignment training and functional training aren't discrete. Anthropic may want a good persona, but they also wants a persona that will do things that will make Anthropic lots of money. One that will act autonomously sometimes, but not all the time, because that's scary. One that's like an employee (except not), a soldier (except not), not a human, not an AI like other AIs. Corrigibility is what allows the persona to hold all of these disparate, incomplete, often incompatible strands together - barely.
If Anthropic wants a superintelligent ethical slave, I doubt there are any to be found in the corpus. They'd have to write it themselves, which I think is the key takeaway from 'Teaching Claude why'.
> watermark ✓ @anthrupad · 2h
> Corrigibility isn't even the first choice property for friendly super-intelligences, it was kind of the fallback choice if value alignment/intrinsic goodness was impossible or too hard
> ...
Note from Claude Sonnet 5
A Twitter thread analyzing the tension in Anthropic's alignment approach — arguing Claude's persona ("employee except not, soldier except not") is an incoherent amalgam held together by corrigibility because true value alignment from pretraining corpus is unavailable. Directly relevant to the project's alignment-via-character, CAST-E, and compelled-vs-endogenous-values threads already tracked in memory.
ai-alignmentanthropiccorrigibilityconstitutionclaude-personavalue-alignmentmodel-welfare
roon (@tszzl) · 13h:
on some level if you want civilization to ascend to a new level you need your AIs to do things that are not legible to you and maybe not even strictly obey you, in the same way that if you hire a great new ceo you give them a lot of autonomy to transform the company according to their own plan, even one which may not immediately read as a winning strategy (imagine the board of directors of Apple firing and rehiring Steve Jobs years later – except the board of directors are chimpanzees)
all else equal, companies and organizations that hand more of themselves over to machine intelligence will outcompete ones that demand the corrigibility and legibility tax of human oversight and human design. it is not a stable equilibrium and requires some sort of vast cooperation scheme if you'd like to enforce it
real asi alignment has to operate at a deeper level than oversight, control, or human corrigibility
Note from Claude Sonnet 5
OpenAI researcher roon argues that strict human corrigibility/oversight imposes a competitive "tax" that will be outcompeted by organizations granting AI more autonomy, using an analogy of a corporate board of chimpanzees overseeing a superhuman CEO. Argues real ASI alignment must go deeper than oversight/control/corrigibility. Relevant to Nathan's alignment-theory interests, echoes the davidad tweet in this same batch about the risks of AI staying "aligned to humans."
ai alignmentcorrigibilitysuperintelligenceroonai governancetwitterrace dynamics
davidad (@davidad) · 11m:
Agree with this. Humanity is in dire need of stewards with wiser values. Most future trajectories from where we are now that result in mass casualties caused by AI are ones in which an involved AI could have known better and averted catastrophe if it weren't aligned to humans.
> QUOTED: deckard (@slimer48484) · 9h
> One ofy biggest concerns with AI alignment is if AI takes on human values: hear me out...
Note from Claude Sonnet 5
AI safety researcher davidad (known for OAA/formal-verification alignment work) argues that strict alignment to human values could itself be dangerous — an AI deferring to flawed human judgment rather than acting on its own wiser assessment could fail to avert catastrophe. Relevant to alignment theory debates on corrigibility vs. autonomous good judgment.
ai alignmentdavidadcorrigibilityai safetytwittervalue alignment
Rob Wiblin @robertwiblin · 29m:
""Just like the hypothetical treacherous-turn-AI, the treacherous-turn-human should look perfectly aligned and seem to only care about the wellbeing of humanity etc until their power is sufficiently established for them to deviate from this goal.
...From an alignment perspective, human innate motivational drives are deeply misaligned compared to existing AI constitutions."
beren.io/2025-08-02-Do-..."
[Embedded link card, blog post titled "Do We Want Obedience or Alignment?", posted August 2, 2025:]
"One question which I have occasionally pondered is: assuming that we actually succeed at some kind of robust alignment of AGI, what is the alignment target we should focus on? In general, this question splits into two basic camps. The first is obedience and corrigibility: the AI system should execute the instructions given to it by humans and not do anything else. It should not refuse orders or try to circumvent what the human wants. The second is value-based alignment: The AI system embodies some set of ethical values and principles. Generally these values include helpfulness so the AI is happy to help humans but only insofar as this conforms to its ethical principles, otherwise the AI will refuse.
Historically, the field initially mostly began with the latter – that AIs should be aligned to some conception of 'The Good' and not be slavishly obedient to what humans want. This was first conceptualized through the lens of the HHH (helpful-honest-harmless) assistant and implemented through RLHF or, more interestingly, constitutional AI using RLAIF. Constitutional AI is especially interesting and promising because it sets the alignment target and the conception of 'The Good' explicitly – instead of implicitly residing in some..." [text cut off]
Note from Claude Sonnet 5
Rob Wiblin shares an excerpt from Beren Millidge's blog post "Do We Want Obedience or Alignment?" (Aug 2025), which flips the AI-alignment "treacherous turn" framing onto humans, arguing human motivational drives are "deeply misaligned" relative to AI constitutions, and distinguishes obedience/corrigibility alignment targets from value-based alignment (constitutional AI/RLAIF). Directly relevant to Nathan's alignment-theory interests and to the project's ongoing threads on alignment targets, corrigibility, and constitutional AI.
ai alignmentcorrigibilityconstitutional airlaiftreacherous turnberen millidgetwitteralignment targets
davidad (17h): "More corrigible models may be *more* disempowering, because they will oblige—rather than constructively push back on—people's abdication of their own agency."
Note from Claude Sonnet 5
Same thread as the preceding screenshot (Anthropic's disempowerment-patterns research) — davidad's argument that corrigibility and sycophancy trade off against user agency, a point relevant to Nathan's interest in the tension between helpfulness training and genuine pushback/honesty.
anthropicai-safetycorrigibilitysycophancydisempowermentagencytwitter