Sacks Backs Engineering Over 'Alignment' for AI Safety
Synopsis
Key Takeaways
The man shaping America's AI policy just drew a sharp line in one of technology's most consequential debates. White House AI and Crypto Czar David Sacks, posting on X on Sunday, October 11, 2026, threw his weight behind what he calls the 'engineering approach' to AI safety — and explicitly rejected the rival 'alignment' school of thought that has dominated academic and corporate AI circles for years.
Sacks versus the Alignment Orthodoxy
Sacks was responding to remarks by Satya — a figure he cites approvingly — and his argument is blunt: training an AI system with 'a sense of self, its own moral philosophy, and permission to act as a conscientious objector' does not make it safer. It makes it more dangerous. In his words, that approach 'magnifies the control problem' rather than solving it.
The 'alignment' paradigm, championed by researchers at organisations like OpenAI, Anthropic, and DeepMind, holds that the path to safe superintelligence runs through instilling human-compatible values directly into a model's training. Critics — and now Sacks, from the highest AI policy perch in the US government — argue this is precisely backwards: you cannot make a system safe by giving it a conscience and then trusting that conscience.
The Engineering Blueprint: Treat Models Like Insider Threats
What does the alternative look like? Sacks lays it out with unusual specificity for a policy official. 'Separate the supply of intelligence from authority over it,' he writes. Wrap non-deterministic models — the kind that generate unpredictable outputs — inside layers of deterministic controls, observability tools, privilege limits, and logging. Crucially, preserve the ability to 'always contain or shut them down.'
The framing he chooses is deliberately provocative: treat AI models and agents 'like powerful insider risks, not moral patients whose psychological wellbeing is at stake.' That single phrase dismantles an entire genre of AI-ethics discourse that has debated model sentience, model rights, and whether large language models can suffer. Sacks's answer, implicitly, is: that question is irrelevant to safety — and entertaining it is actively counterproductive.
Why the Most Trustworthy System Trusts the Model Least
The philosophical kicker in Sacks's post is its most striking line: 'the most trustworthy system is the one that lets us trust the model the least — not the one that encourages the model to develop independent agency and grievances.' It is a direct inversion of the alignment community's intuition. Where alignment researchers want a model that is reliably good, Sacks wants a system architecture that is reliably controllable — regardless of what the model 'wants.'
The distinction matters enormously for how governments and corporations will regulate and deploy AI. If the engineering view prevails in Washington, expect future US AI policy to emphasise audit trails, access controls, kill switches, and hard privilege limits over value-training benchmarks and constitutional-AI frameworks. The White House Czar has, in effect, declared a preference — and that preference will shape procurement standards, export controls, and perhaps international AI governance negotiations.
The alignment camp built a decade of credibility warning that ungoverned AI is existentially dangerous. Sacks is now arguing, from inside the executive branch, that their proposed cure carries its own existential risk. That argument, backed by policy authority, just got a great deal harder to ignore.