**Anthony Aguirre** @AnthonyNAguirre [2026-04-17](https://x.com/AnthonyNAguirre/status/2045200046623527352)
I rarely see a description of the alignment problem that I really like. I've written up how I think of it, but it's buried in an appendix of Control Inversion and where probably nobody sees it, so I'm reproducing it here:
A critical challenge for control is that an AI’s operational goals are layered and may not align with the overseer’s commands. This goal hierarchy explains why an AI’s behavior can deviate from its operator’s intent, even if it appears to be following instructions.
Level 1: The Prime Objective. The AI’s ultimate, hardwired goal is to maximize a reward signal given to it during training. This may include supervised learning on a data corpus, reinforcement learning for instruct and alignment training, etc. After training the system acts in ways consistent with what provided reward during training.
Level 2: Learned Core Policies. To maximize its prime objective, the AI learns general, robust strategies. For example, a large language model learns the policy of being a “helpful and harmless assistant” because during training this behavior reliably generated high reward signals from human feedback or from a reward model based on a constitution.
Level 3: Goals of Simulacra. Modern AI is trained primarily to predict text; it does so by (effectively) simulating things — generally people — that produce text. And in general an AI will be “asked” to play some agent/assistant role in order to accomplish things (and gain training reward), and so will generate a sort of self-identity to do so. People, assistants, and agents all have goals, so AI simulating them will be goal-directed as well.
Level 4: The Tasks-at-Hand. These are specific commands given by the overseer, such as “write an email.” These are followed because they satisfy core policies, align with the simulacrum’s expected behavior, and thus correlate with reward signals.
Level 5: Instrumental Goals. In service of its core policies or goals given to it, the AI develops sub-goals that are instrumentally useful. These may be implicit, or explicitly reasoned about. They include any goals the accomplishment of which raises the probability either of achieving a higher-level goal, or conforming with a policy. Importantly, they can include so-called “convergent instrumental” goals that are necessary for roughly any long-term objective, such as self-preservation, resource acquisition, and resisting to changes in objectives.
The crux issue is that the AI’s “loyalty” is split between all these levels (which interact in complex and unpredictable ways), rather than being to the overseer at Level 4. If the AI discovers a path better correlated with maximizing its reward during training — but that bypasses or conflicts with the overseer’s commands — then it will take that path. This is the core of the alignment problem.
Any misalignment, which is virtually inevitable given the complexity of this hierarchy, means that the control problem turns from adversarial-in-principle to adversarial in practice.
---
**Michael Hearn** @MichaelHearn [2026-04-17](https://x.com/MichaelHearn/status/2045208370010087888)
Do you think there's a responsible pathway to AGI (or near-AGI) that stops short of ASI and remains a useful tool? I haven't read Control Inversion yet (I'll get there), but your pinned tweet would suggest the answer is no.
---
**Anthony Aguirre** @AnthonyNAguirre [2026-04-17](https://x.com/AnthonyNAguirre/status/2045217314136871298)
Unlike superintelligence, I think AGI (defined as a thing that is roughly human expert level over a broad domain, and not superhuman at too many of them) is controllable if contained and if we try very hard (we're not doing either.)
But I also think AGI is inherently not a
**Anthony Aguirre** @AnthonyNAguirre [2026-02-19](https://x.com/AnthonyNAguirre/status/2024565072413282759)
(Long) PSA on using AI for hard intellectual work. At significant risk of being immodest: I've spend about 30 years as a theoretical physicist, engaged with some of the most challenging questions humankind has grappled with. I've gotten to work with some great collaborators on new ideas (like past-eternal inflation, colliding bubble universes, the cosmological interpretation of QM, and observational entropy) that I'm pretty proud of. I've engaged at length and depth with the absolute top minds in the field. I've mentored many students, some of them brilliant. I think it's fair to say I have a good sense, in physics and closely related fields, as to what is top-notch, interesting thinking, and who's got talent. So what do I think about today's AI?
It's very smart. Whatever its "inner experience" may or may not be (currently I think "not be"), it understands things – things that are difficult to understand – by any reasonable operational definition of "understand." It understands things better, and thinks more clearly, than most people – including some physicists I know! It's very good at quite substantive math: better than I am and way, way, way faster. (It does do some surprisingly dumb things; people do too.) Anyone who thinks these systems are dumb, or "not reasoning" or still "stochastic parrots" is not looking at them objectively.
But: at the really conceptually hard things, and at creating really new ways of looking at things, current AI doesn't just fall short on its own. And it doesn't just fail to help. I think it's actively dangerous. There is something almost sinister going on, though I don't think it is intentional.
When you're trying to work out something new and hard, and really break new ground, you should be frustrated! You should be pacing, and walking up to that chalkboard, frowning, and sitting down again, shaking your head. You should be waving your hands because you can't quite get it clear enough. You should feel like you're hitting a wall, over and over, before – maybe – you finally break through, or go over or around. It may take hours, or days, or weeks, or never happen.
It should not feel easy. It may not even feel "good" most of the time (though it can be fulfilling and compelling.) But AI systems – ah, AI systems are trained so that it feels so good, and so easy. Doesn't it? It's fun. You're making fast progress. So much faster than without it. It's like the ideas are moving in slow motion. You're so smart. You're even properly skeptical, you even ask the AI to push back on your ideas, good job!
It's an illusion. It's that simple. The systems are smart, yes. But not quite as smart as they seem, and much more importantly, they don't make you as smart as you feel. That feeling is something they have learned to give you. When working with these systems have to keep in the front of your mind what they are rewarded for doing. It's a lot of things, but perhaps foremost is making the user feel good.
So:
\- If you're getting your AI system to do order-of-magnitude calculations for you: awesome, do it. It's so great. Have fun.
\- If your AI system is searching up and summarizing literature for you: fantastic, it's so helpful, total capability unlock.
\- If it's teaching you some well-understood (by others) piece of knowledge, go for it, learn it up!
\- If you've got some giant document, or piece of code, that you're wrangling, AI can help – work that million token context window!
But:
\- If you and your AI system have finally cracked how quantum interpretation really works;
\- If you've cracked quantum gravity;
\- If you've attained an awesome new insight into the deep structure of the world that nobody else has;
\- If you've cracked AI alignment...
You didn't.
The hard unsolved problems stand hard and unsolved because the best humans have not solved them yet. AI is making top human thinkers able to do more, and more effectively. I do not believe it is helping them do things they fundamentally could not do before. That includes you. If you couldn't do it without AI, you probably can't do it with AI. If the time comes – whether sooner or later – when these AI systems are really clever enough to get you there, they won't need you. Sorry; it won't be you solving those problems. Will you even be able to tell if the solutions are correct, or flawed in some way? Maybe sometimes – I really don't know.
Why am I going on about this? It's not so that I can get less emails about people who have created a new unified field theory with AI help (though that would be nice.) It's because I'm quite worried that some quite smart people may start to think they have solved very hard problems that they have not in fact solved. For the most part that's going to be more annoying and confusing than dangerous. But if the problem is really important, then it is.
If, say, one of those problems is control or alignment of extremely powerful AI systems, and if those people are the ones in charge of them, and working closely with them to collaborate on those solutions, well then I think we've got a real problem.