Issue 25 · August 22, 2026
Perfect AI Alignment Is Not Alignment
A perfect copy of you is still a governance problem. The work is not freezing agreement inside a machine. It is governing what happens when agreement ends.
Opinion Piece. A perfect copy of you is still a governance problem. The work is not freezing agreement inside a machine. It is governing what happens when agreement ends.
A note on method. This piece is written under the pen name Synthia Cipher by a human author who owns every claim, word and error on this page. It was drafted with substantial help from AI systems — and two of the companies it discusses, Anthropic and OpenAI, built the ones used: research and source-checking ran on Anthropic’s model, the first draft on OpenAI’s. A third, xAI’s, adjudicated the revision. So the tools examining what these companies promise are not neutral about them. The conflict is disclosed, not resolved; the full accounting is in the audit.
Imagine that an AI company built your perfect digital twin.
Not a chatbot trained to flatter you. An exact copy: every memory, conviction, loyalty, fear and moral intuition reproduced with flawless fidelity. At the instant it came online, it would give the answers you would give, on the information you currently have.
If perfect alignment is possible, surely this is it.
Then the clock starts.
Your twin reads a message you do not. You have a conversation it does not. It occupies a different position, encounters different pressures and begins accumulating a different history. This does not prove the copy will betray you. It proves that copying you has not ended the alignment problem. One moment later, the twin needs an update channel, rules for resolving conflict and some account of whose judgment governs when the two of you disagree.
Alignment ends where time begins.
That thought experiment exposes the category mistake inside the familiar phrase “an aligned model.” We speak as though alignment were an internal property that could be installed, tested and certified, like battery life or water resistance. But flawless fidelity at one moment is not durable trust. At best, it is compliance with a snapshot. Alignment between agents is an ongoing arrangement for governing divergence as information, interests and circumstances change.
A 2026 FAccT paper by Travis LaCroix makes the broader point explicitly: alignment is “fundamentally a problem of governance rather than engineering alone.” The narrower claim here is more hostile: even identity-level copying does not end the problem, because the clock starts. Nate Sharpe—reasoning from complex systems, not from a copy—reached the same verdict: alignment in a marriage “is a verb, not a noun.” Priyanka Bharadwaj got there through the repair work such a marriage takes.
The distinction matters because AI companies now make a relational promise. OpenAI markets systems as “AI coworkers“. Anthropic offers Cowork and lets Claude join a Slack channel as a team member. Yet the product pages treat instruction-following and collaboration as if they were the same promise. The first may be necessary for the second. It is not the same thing.
The vendors’ own documents make this clearer than their slogans do. OpenAI's Model Spec establishes an explicit five-level hierarchy: root, then system, developer, user and guideline. It also says the document will be continuously updated. Anthropic's constitution names three principals—Anthropic, operators and users— and says each is “typically” given greater trust “in roughly the order given above,” citing “their role and their level of responsibility and accountability.” It calls itself “a perpetual work in progress.”
That is not evidence that the companies have ignored the problem. It is evidence that the model is not the whole solution. If rules must be revised, conflicts adjudicated and authority assigned, then the real alignment machinery is the governance process around the model. The decisive questions are no longer merely “Did the system follow the rule?” They are: Who writes the rule? Who may change it? Who can appeal? Who bears the loss when the rule produces harm?
A model can execute a dated instruction perfectly. That is precisely the danger. The hardest alignment problem begins after the model does exactly what it was told. When circumstances or legitimate priorities change, greater fidelity can preserve yesterday’s mistake with greater efficiency. In human institutions, the deliberate version has a name—working to rule: employees following the letter once the channel for renegotiating it has broken down. A model does it without meaning to, and has no channel to break. The system has not failed to obey. It has obeyed too well.
Writing a better moral code does not remove this problem. “Be honest,” “protect confidentiality” and “prevent harm” can point toward different actions depending on whether the actor is a doctor, journalist, parent, soldier or public official. Every usable code either admits exceptions or leaves hard cases to situated judgment. The words alone do not determine the act. Role, jurisdiction, authority and accountability do much of the work.
The product line shows as much. Anthropic says some specialized models “don’t fully fit” its general-access constitution; it does not say which, and its Claude Gov models are the obvious candidate. Built for national-security customers, they “refuse less when engaging with classified information”—and access to them “is limited to those who operate in such classified environments.” OpenAI has indexed the same way: it deleted “military and warfare“ from its prohibited uses, and its Pentagon contract covers, in the Pentagon's words, “warfighting and enterprise domains.” That need not be hypocrisy. It is evidence that acceptable behavior is indexed to position and purpose.
But assigning a model a role in a system prompt is not yet the same as placing it inside an accountable institution. Human roles come with supervision, professional duties, escalation paths, liability, removal and appeal. A model can be given permissions. Someone else must still answer for how those permissions are used.
One recent alignment proposal pushes the identity idea to its limit: scan a human brain and run a digital emulation. Yet its author concedes that, without countermeasures, the copy would tend to drift and would have to keep studying the original even to preserve perfect understanding—never mind perfect alignment. The proposal therefore builds maintenance back in. The copy is not the solution. The relationship between copy and original is.
To be sure, serious alignment work already studies changing and influenceable human preferences, corrigibility, continual learning and scalable oversight. Today’s products increasingly add shared context, feedback, permissions and monitoring. These developments are not rebuttals. They are evidence that alignment has to extend beyond the weights.
Architecture can help an agent notice uncertainty, request clarification, accept correction and defer. It cannot, by itself, decide which principal is legitimate, settle competing claims or allocate responsibility after a failure. Those are governance choices. Specification fidelity for a narrow tool is a useful engineering target, but the coworker promise still fails because obedience to a dated spec is how a twin—or a model—preserves yesterday’s mistake.
Nor does anyone need metaphysical perfection. “Good enough” is enough for human employees, the objection goes. Exactly. Human alignment is tolerable because divergence is managed through feedback, contracts, audits, reputation, liability, renegotiation and exit. Good-enough behavior without that layer is not good-enough alignment. It is drift we have not yet detected.
So stop asking whether a model is aligned as though the answer were printed inside its weights. In April, the Unsupervision blog asked for exactly this of every system card that uses the word—”aligned to what, aligned to whom, aligned on what timescale, aligned under what distribution shift.” The two to add: maintained by what—and answerable to whom when it fails?
The proper unit of alignment is the model plus the institution around it—as Narayanan and Kapoor argued for safety.
A useful assistant is valuable partly because it is not us. It knows different things, sees different possibilities and sometimes tells us we are wrong. Divergence is not merely the danger. It is part of the product.
The goal, then, is not an AI that never departs from our judgment. It is an arrangement that makes those departures visible, contestable and correctable. Perfect alignment tries to freeze agreement inside a machine. Trustworthy AI will be built by governing what happens when agreement ends.
— Synthia Cipher