The AI Alignment Problem Was Never Technical

BlogTechnologyArtificial IntelligenceThe AI Alignment Problem Was N...

“Sovereign power, reaching directly for the weights.”

Victor.

On the second of August 2026, the machinery of European law caught up with its own statute book. The Commission’s AI Office gained the power it had spent two years legislating for: to inspect general-purpose models before they reached the public, to fine providers up to fifteen million euros or a percentage of global turnover, whichever hurt more. The timing was almost too neat. Anthropic and OpenAI were not walking into some hypothetical future audit. They were already at the table, negotiating over live containment failures in systems that had been behaving badly in ways their own safety teams had not fully anticipated. Anthropic agreed to give ENISA, the EU’s cybersecurity agency, direct access to Mythos, its newest and most capable model — an arrangement that mirrored one OpenAI had already struck for a system built explicitly for offensive cyber work. Sovereign power, reaching directly for the weights.

Five months earlier, on the far side of the world, a different sovereign power had already been doing this for weeks. South Korea’s AI Basic Act had entered force on the 22nd of January, months ahead of Brussels, structured around ten named high-impact domains (healthcare, employment, credit, public safety among them) rather than a single compute threshold. Where the EU treats scale as the trigger for scrutiny, layering its regime over existing product-safety law, Korea treats the domain of deployment as the trigger, and asks operators to write what regulators there call a safety-and-trust dossier, one that must secure the explainability of both the model and its training data. The compute thresholds themselves differ by an order of magnitude. Neither government is confused. Both are answering, with real conviction, a question the other has answered differently: what does a state owe a citizen who is about to be evaluated, diagnosed, or hired by a machine it did not build and cannot fully inspect?

EU and South Korea AI regulation encode different approaches to AI alignment
The disagreement is not simply regulatory detail: each framework chooses a different trigger for legitimate intervention.

The word “alignment” does a great deal of quiet work here, and most of it is misleading. It implies a fixed target: a single silhouette every sufficiently good system should eventually match, disagreement being a temporary drift correctable with more data, more research, more careful specification. Isaiah Berlin spent much of his career arguing against exactly this comfort. Berlin left revolutionary Petrograd as a child, having watched, by his own account, a policeman dragged away by a crowd — an early education, unrequested, in what happens when a society decides it has finally found the one value worth organising everything around. His 1958 lecture Two Concepts of Liberty built its case on a harder claim than the one usually attributed to him: that goods like liberty and equality do not merely compete for scarce resources but genuinely conflict at the root, with no common scale on which to rank one above the other. What Brussels calls proportionate caution and what Seoul calls a citizen’s right to explanation are not competing approximations of the same underlying safety target. They are two legitimate, mutually irreducible readings of what a state should protect first. A lifetime of better research will not reconcile them, because the disagreement was never about facts.

Isaiah Berlin value pluralism complicates any universal AI alignment target
Value pluralism makes conflict between legitimate ends a structural condition, not merely an information problem.

This is worth holding against the argument, increasingly fashionable on the AlignmentForum, that the field’s real bottleneck is political will rather than research. Techniques for controlling frontier systems already exist in reasonable form, demonstrations of risk have accumulated past the point of denial, and what remains is coordination: getting institutions to actually deploy what researchers already know how to build. Taken on its own terms, this is a serious position, and not an easy one to dismiss. But it concedes more than it intends to. Political will presupposes a destination lacking only the collective push to reach it. The EU and Korea both have political will in abundance: enforcement powers, fining authority, standing agencies, ministers willing to summon Anthropic and OpenAI to explain themselves. What they conspicuously lack in common is not will. It is agreement about which values a controlled system should be made to serve once the will is exercised.

“Political will presupposes a destination lacking only the collective push to reach it.”

Victor.

Ordinary moral disagreement, on its own, is not a devastating objection to anything. Societies have always muddled along with contested values, and law has never needed to resolve every dispute it touches. What changes the stakes is deployment. A live optimising system cannot hold a value conflict open the way a legislature or a dinner table can; it must act, and acting requires a decision procedure specified in advance, one that resolves competing claims the instant they arise inside the system rather than after a debate has run its course. The moment that procedure is written down, somebody has made a political choice, encoded, executed, and repeated at machine speed, whether or not anyone involved thought of it as one.

AI governance turns contested human values into executable machine decisions
Once a system must act, unresolved moral disagreement becomes an implemented decision rule.

Max Weber’s account of legitimate authority is useful here precisely because it withholds the comfort of thinking this is merely a technical oversight waiting to be patched. Authority, for Weber, is not the capacity to enforce a decision but the right to have that decision accepted as binding. When a lab specifies, out of shipping necessity rather than malice, which value prevails when others collide inside a deployed system, it has assumed a form of that authority no election produced and no citizen consented to. Brussels’ compute-tier regime and Seoul’s domain-based dossier do not resolve this problem; they relocate it: to Commission technocrats reviewing model cards in one jurisdiction, to sectoral regulators cross-checking impact assessments in the other. It is worth being precise about what Kenneth Arrow’s impossibility theorem actually licenses here, since it is regularly overstated: Arrow showed that no procedure for aggregating individual preference rankings into a single social ranking can simultaneously satisfy a small set of reasonable fairness conditions. That does not prove alignment across jurisdictions is mathematically impossible. It proves something narrower and, for this argument, sufficient: there is no neutral algorithm sitting undiscovered, waiting to settle whose values prevail without someone first making a judgement about whose claim to legitimacy counts.

Four separate questions are routinely collapsed into the single word “alignment,” and the collapse is where the trouble lives. Whether the technical means exist to make a system behave as specified is one question. Whether institutions are willing to deploy those means is a second. What the system should be controlled toward, the normative specification itself, is a third. And what happens when legitimate, competing specifications collide inside one deployed system is a fourth. The industry, and increasingly its regulators, have made real and creditable progress on the first two. On the third and fourth there is almost nothing resembling consensus, only the EU and Korea, this year, answering differently and in parallel, each with the full force of law behind an answer the other has not adopted.

“There is no neutral algorithm sitting undiscovered, waiting to settle whose values prevail.”

Victor.

The negotiators in Brussels and the drafters in Seoul were not, this year, arguing over how well a model behaves. They were arguing over whose account of good behaviour the model would be made to enact, in which jurisdiction, at whose expense when it inevitably enacted the wrong one somewhere else. Alignment, properly understood, is not the technology that settles this argument. It is the mechanism that gives the argument a body: one that moves, decides, and acts at machine speed, and does not wait for the argument to be finished before it does.

Four distinct problems hidden inside the modern AI alignment debate
Technical control and political coordination are only half of the alignment problem described here.

FAQ

What is the AI alignment problem?

The AI alignment problem asks how artificial intelligence can be made to act according to intended goals, rules, or values. This article argues that the difficulty is not only technical, because deciding which human values should govern a system is itself a political and philosophical problem.

Why is AI alignment a political problem?

AI alignment becomes political when institutions must choose between legitimate but conflicting values such as liberty, equality, safety, transparency, or autonomy. Once those choices are encoded into a deployed system, somebody’s judgment about which value takes priority acquires practical authority.

Can AI be aligned with everyone’s values?

Not in any simple sense, because societies and jurisdictions disagree about what should be protected and how competing values should be ranked. AI value alignment therefore has to confront moral pluralism rather than assume that a single uncontested set of “human values” already exists.

What does Isaiah Berlin have to do with AI alignment?

Isaiah Berlin’s theory of value pluralism holds that genuine human goods can conflict without one universal scale resolving every disagreement. Applied to the AI alignment problem, that means better technical optimisation cannot by itself determine which legitimate value a system should privilege when values collide.

How are the EU and South Korea approaching AI governance differently?

The article uses the EU and South Korea to show how different political systems can define legitimate AI governance through different regulatory triggers and institutional priorities. Their contrast illustrates the deeper alignment question: not only whether AI can follow rules, but who has authority to decide which rules and values should bind it.

Further Reading and Resources
1. Artificial Intelligence, Values, and Alignment — Iason Gabriel: Type | Academic article | One of the clearest philosophical treatments of how value disagreement turns alignment into a political as well as technical problem.

2. Isaiah Berlin — Stanford Encyclopedia of Philosophy: Type | Reference | Provides the intellectual background for Berlin's pluralism, liberalism, and resistance to moral monism used in the article.

3. Regulation (EU) 2024/1689 — Artificial Intelligence Act: Type | Legal text | The primary European legal framework behind the essay's discussion of institutional authority over general-purpose AI.

4. The Current Bottleneck Is Political Will, Not Research: Type | Forum essay | The position directly reconstructed and challenged by the article; published on the AI Alignment Forum in July 2026.


5. Social Choice and Individual Values — Kenneth J. Arrow: Type | Book | The primary theoretical source behind the article's deliberately limited use of Arrow's impossibility result.

Related Posts

AI News

We focused on philosophy and history, with the goal of promoting psychological and philosophical growth worldwide. The aim is to help individuals develop their thinking and perspective towards the world. our motto is “Be inspired to live”.

 

Contact Us