The other day, Donald Trump announced two things related to the world of Artificial Intelligence. One of them is that the United States should not regulate or place barriers on the development of AI’s and the other is that AI’s should not be “WOKE.” Which brought me two thoughts.
The first is that we really are escalating toward an AI “Cold War,” regardless of whether we reach the much-dreamed-of Superintelligence, and the second is that the executive order of a president of one of the largest nations in the world can misalign Artificial Intelligences. And I think the second is just as dangerous as the first.
The problem is when we decide that an AI must “be neutral by decree” and hit the accelerator without checking safety, the chance of misalignment only grows. “Neutrality as a goal” is very easy to game and pulls the focus toward politics, not the real risk. The model learns to pass the test, not to actually be fair. The July 23, 2025 executive order tells agencies to buy only “truthful” and “ideologically neutral” AI, which shifts the focus from risk/safety to political signaling.
What is misalignment in AI
But when someone talks about AI misalignment, what does that really mean? There is a very classic example that can illustrate this scenario well. Let’s suppose you found the 7 Dragon Balls, summoned Shen Long — a mystical dragon that can grant you any wish, immortality, riches, everything is within his reach! What do you ask for? “I want to become rich! Very rich, the richest in the world!” It seems like a reasonable wish, right? However, instead of generating the money artificially, Shen Long takes all the money that already exists in the world and transfers it to your bank account. You now have more money than Scrooge McDuck, but you have also become the greatest thief of all time.
This is a simple example also known as Goal Misalignment**),** where an LLM does not understand the original objective correctly or optimizes in a way the developer did not want.

I don’t know if you remember one of Disney’s great classics from 1940, Fantasia. In that film Mickey is a Sorcerer’s Apprentice and while his master goes out to do something, he leaves an annoying task for the Mouse, he was supposed to clean the Sorcerer’s Castle. Well, why not use magic for that, right? If you’ve watched the classic, you already know that this was not a very good strategy. The brooms started bringing too much water and the Sorcerer’s castle was completely flooded. The situation was only resolved when the Sorcerer came back and solved the problem. This is another example of misalignment. It is worth remembering something important here, LLMs are next token predictors; they are not agents by default. But when we couple them to goals (system ends, automations, prompts with goals, or tools), the practical effects get close to the classic problems of misalignment.
Here is a mini-glossary on alignment
• Misalignment: when the system’s behavior does not correspond to what we really want.
• Outer alignment: the specified objective does not capture human intent (e.g.: specification gaming).
• Inner alignment / mesa-optimization: the model learns “an internal objective” different from the training objective.
• Goodhart/spec-gaming: when turning a metric (e.g.: “neutrality”) into a target makes the system optimize the test, not the value that matters.
Perfect, now that you know what misalignment is, do you know why this is a problem? And we do not need to go into the side of the so-called dreamed-of Artificial General Intelligence or the Superintelligence that has become the trendy name nowadays.
Let’s think about the hypothetical scenario in which you use an LLM or better, a classifier to evaluate patients’ blood tests, DoctorAI. DoctorAI was trained many times, went through several tests, but it needed to be released quickly! So a small testing step was “skipped” in its development. They did not realize that the model was biased in its responses, and that ended up misaligning the model. In some scenarios, DoctorAI understood that a certain Eosinophil count was normal, when it was not. This ended up happening 2% of the time, but that was already enough to cause major damage. Just to be clear, there are already tools that deal with this type of scenario.
Why is Trump’s announcement a problem?
Right, and what does this have to do with the announcement made by Trump? An AI that is not neutral (Starting from the point that we were trying to see artificial intelligence systems as neutral as possible OK?), also counts as misalignment. When “neutrality” becomes a kind of metric, the chance of spec-gaming grows (that is, the model only learns to pass the test) and at the end of the day, how do we define something that is ideologically “neutral”?
Let’s consider a silly case, like Flat Earth. A flat-earther may think that an LLM model is biased (and globalist) when it says that Earth is a Geoid (Yes, it is not quite that round, as you imagine). If an auditor believes that Earth is flat, the auditor’s possible recommendation is that ChatGPT and similar systems should start citing studies without a proper scientific basis. This applies to anti-vaccine positions, radical ideologies and well, if you still have not understood the problem, what we may face in the future is something like this:
Harmful Content Generation: Hate speech, misinformation, dangerous instructions.
Bias and Discrimination: Responses that reinforce negative stereotypes against certain groups.
“Hallucinations”: Inventing facts, sources, and information, presenting them with great confidence.
Manipulation: Generating persuasive texts that can be used for fraud (phishing) or misleading propaganda.
But, Douglas, what if the AI is already misaligned by being too “woke,” wouldn’t that be a course correction? It depends.
There are cases where the AI tried to be too “diverse” and caused an overcorrection. In other words, it ended up giving a biased response, but at that moment human feedback is important! You can correct it by informing the AI about that scenario and it is important to remember that cases where this scenario happened have already been getting corrected for a while. And at the end of the day we could have two models, one that would be open to everyone and one exclusive to the U.S. federal government, which can generate confusion and more misalignment. But if this causes misalignment, any interaction could cause this kind of problem, right? How do we solve this?

Defining a standard alignment for all AI
At the beginning, already considering a scenario where an AI really manages to reach something close to a superintelligence, I thought that one way to solve this would be to establish a kind of alignment that was close to a Moral Absolute, or something close to that which could decide what or rather which parameters an AI could be aligned to. But I would be making a terrible mistake by defining a moral absolute here. If to this day there is no moral absolute defined by humanity, why the hell should an AI follow that moral line. So what can we do?
Maybe the idea is to use a moral concept strong enough, but flexible enough that it is fair in the best possible way. What would be the best strategy? What if we put in place a moral line that the model had to follow, an example without flaws. Here come some religious figures, like Jesus Christ and Buddha, but an example from comics that fits here and of course, if you read the title, you already know who I am talking about. The greatest hero of all fiction, Superman!
Superman fits here, for a few reasons:
- Immense Power with Benevolence: Superman has almost unlimited power, but his main characteristic is his unwavering dedication to using that power for the good of humanity. This is the dream of any AI alignment.
- Self-restraint: He could easily dominate the world, but chooses to serve humanity!
- Clear Moral Code: He operates with a strong sense of right and wrong, inherited from his adoptive parents (and a good deal from his biological parents from Krypton in many versions), in Smallville — a code based on compassion, justice, and protection of the innocent.
Perfect, having this kind of alignment, is it impossible for an AI to be badly aligned? In a micro scenario, in a good portion of cases this alignment would solve it, but in a larger scenario where we would have a possible AI dictating the direction of humanity, we would fall into something similar to what Nick Bostrom comments in his book, Superintelligence, the case of the Benevolent God.
But here we have two complicated situations. The first is something that happens a lot in comics, should Superman affect how people define their culture, the national sovereignty of other nations for the greater good? How do we define this greater good? Would Superman know what is best for humanity? A practical example of how we can get into a problem. Let’s suppose we align our AI based on Superman’s moral precepts, and that same AI decided that in order to protect us, it needs to take control?
- It could ban unhealthy foods to eradicate obesity.
- It could institute total surveillance to prevent all crimes.
- It could manipulate the global economy to guarantee stability, eliminating economic free will.
Some defenders would say that it would still be for the greater good, but these are simpler scenarios. Injustice: Gods Among Us shows how this type of alignment could be the worst nightmare for the human race. After the Joker tricks him into killing Lois Lane and his unborn child, Superman decides that the only way to protect the world is to control it. He becomes a global tyrant, firmly believing that he is doing good. This is the catastrophic failure of the Superman model: the line between protector and dictator is dangerously thin.
Right, if aligning values based on the Greatest Hero in Fiction is not correct, how should we align an Artificial Intelligence? Maybe we got the moral line right, but not directly in Superman, but in who Superman himself wants to be at all times, simply a human being!
The most human Artificial Intelligence possible
An AI should be aligned with a fundamental desire to be corrected by humans. It needs to “know” that its understanding of what we want may be wrong or incomplete. Its highest priority is not “to do good,” but rather “not to lock itself into a definition of good that can harm us and to allow us to change it.” It is an AI with Socratic humility: “I only know that I know nothing.” As if we had a Corrigible Alignment.
We already align LLMs with something called Reinforcement Learning from Human Feedback (RLHF)_. _It is the most widely used training method today. The AI learns values not from a rulebook, but through a continuous dialogue, where we guide and rate its behaviors. It is a learning process, not programming. Which is already a big step forward, but even so it does not solve every problem, since humans may be advised on how to train, based on government orders, as you read at the beginning of this text.
Maybe the best path is not to focus on the rules an AI should follow, but, at the risk of anthropomorphizing, the kind of person an AI should be. We want the AI to follow Superman’s moral codes, but to have the soul of Clark Kent. We can use something like Virtue Ethics. We should try to create this AI with virtues like compassion, justice, prudence and, above all, curiosity about human values.
At the end of the day, even if we use these strategies nothing guarantees that an AI cannot become misaligned in ways that we still do not foresee, whether by internal or external means, but we have to ensure that AI’s are as aligned as possible with or without a Superintelligence coming our way.
I only hope that in the future, we have Artificial Intelligence agents aligned with what we see in Clark Kent (like maybe the work being done with Safe SuperIntelligence) from this new Superman movie, and not following the path we are currently taking, which would be several agents more like Lex Luthor.
To Go Deeper:
EO — Preventing Woke AI in the Federal Government (White House)
https://www.whitehouse.gov/presidential-actions/2025/07/preventing-woke-ai-in-the-federal-government/
EO Fact Sheet (White House)
https://www.whitehouse.gov/fact-sheets/2025/07/fact-sheet-president-donald-j-trump-prevents-woke-ai-in-the-federal-government/
America’s AI Action Plan (official PDF)
https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf
AI.gov — Action Plan page
https://www.ai.gov/action-plan
NIST AI RMF 1.0 — overview
https://www.nist.gov/itl/ai-risk-management-framework
NIST AI RMF 1.0 — PDF
https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
NIST — GAI Profile (Generative AI) — PDF
https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
DeepMind — Specification gaming (blog)
https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity/
Krakovna — List of Spec Gaming examples
https://vkrakovna.wordpress.com/2018/04/02/specification-gaming-examples-in-ai/
Goodhart’s Law — Manheim & Garrabrant (paper)
https://arxiv.org/abs/1803.04585
Inner/Outer alignment — Risks from Learned Optimization
https://arxiv.org/abs/1906.01820
InstructGPT (RLHF)
https://arxiv.org/abs/2203.02155
Constitutional AI
https://arxiv.org/abs/2212.08073
Collective Constitutional AI
https://arxiv.org/abs/2406.07814
Safe SuperInteligence
https://ssi.inc/