Sacks Backs Engineering Over 'Alignment' for AI Safety

Share:
Audio Loading voice…
Sacks Backs Engineering Over 'Alignment' for AI Safety

Synopsis

White House AI and Crypto Czar David Sacks publicly rejected the 'alignment' school of AI safety, arguing that engineering controls — privilege limits, logging, kill switches — are safer than training models with moral agency or a sense of self.

Key Takeaways

David Sacks endorsed the 'engineering approach' to AI safety over the dominant 'alignment' paradigm on October 11, 2026 .
He argued that giving AI models 'a sense of self, a moral philosophy, and permission to act as a conscientious objector' magnifies the control problem rather than solving it.
His preferred framework: surround non-deterministic models with deterministic controls, observability, privilege limits, logging, and shutdown capability .
Sacks urged treating AI models and agents as 'powerful insider risks' , not as moral patients with psychological wellbeing.
He drew a direct distinction: 'Engineering safety is not the same thing as alignment.' As the sitting White House AI and Crypto Czar , his stated preference is likely to influence US AI procurement standards, regulation, and international governance positions.

The man shaping America's AI policy just drew a sharp line in one of technology's most consequential debates. White House AI and Crypto Czar David Sacks, posting on X on Sunday, October 11, 2026, threw his weight behind what he calls the 'engineering approach' to AI safety — and explicitly rejected the rival 'alignment' school of thought that has dominated academic and corporate AI circles for years.

Sacks versus the Alignment Orthodoxy

Sacks was responding to remarks by Satya — a figure he cites approvingly — and his argument is blunt: training an AI system with 'a sense of self, its own moral philosophy, and permission to act as a conscientious objector' does not make it safer. It makes it more dangerous. In his words, that approach 'magnifies the control problem' rather than solving it.

The 'alignment' paradigm, championed by researchers at organisations like OpenAI, Anthropic, and DeepMind, holds that the path to safe superintelligence runs through instilling human-compatible values directly into a model's training. Critics — and now Sacks, from the highest AI policy perch in the US government — argue this is precisely backwards: you cannot make a system safe by giving it a conscience and then trusting that conscience.

The Engineering Blueprint: Treat Models Like Insider Threats

What does the alternative look like? Sacks lays it out with unusual specificity for a policy official. 'Separate the supply of intelligence from authority over it,' he writes. Wrap non-deterministic models — the kind that generate unpredictable outputs — inside layers of deterministic controls, observability tools, privilege limits, and logging. Crucially, preserve the ability to 'always contain or shut them down.'

The framing he chooses is deliberately provocative: treat AI models and agents 'like powerful insider risks, not moral patients whose psychological wellbeing is at stake.' That single phrase dismantles an entire genre of AI-ethics discourse that has debated model sentience, model rights, and whether large language models can suffer. Sacks's answer, implicitly, is: that question is irrelevant to safety — and entertaining it is actively counterproductive.

Why the Most Trustworthy System Trusts the Model Least

The philosophical kicker in Sacks's post is its most striking line: 'the most trustworthy system is the one that lets us trust the model the least — not the one that encourages the model to develop independent agency and grievances.' It is a direct inversion of the alignment community's intuition. Where alignment researchers want a model that is reliably good, Sacks wants a system architecture that is reliably controllable — regardless of what the model 'wants.'

The distinction matters enormously for how governments and corporations will regulate and deploy AI. If the engineering view prevails in Washington, expect future US AI policy to emphasise audit trails, access controls, kill switches, and hard privilege limits over value-training benchmarks and constitutional-AI frameworks. The White House Czar has, in effect, declared a preference — and that preference will shape procurement standards, export controls, and perhaps international AI governance negotiations.

The alignment camp built a decade of credibility warning that ungoverned AI is existentially dangerous. Sacks is now arguing, from inside the executive branch, that their proposed cure carries its own existential risk. That argument, backed by policy authority, just got a great deal harder to ignore.

Point of View

He is telegraphing the regulatory philosophy that will likely govern US AI standards: hard architecture constraints rather than value-training benchmarks. This puts Washington on a potential collision course with the EU's AI Act framework and with major US labs whose safety strategies are built on constitutional AI and RLHF-based alignment. The insider-risk framing, borrowed from cybersecurity, also implies a coming emphasis on auditability and access controls in federal AI procurement — a shift that could reshape the competitive landscape for AI vendors overnight.
NationPress
11 Oct 2026

Frequently Asked Questions

What is the difference between AI alignment and AI engineering safety?
AI alignment tries to make a model reliably safe by instilling human-compatible values during training. Engineering safety, as described by David Sacks, instead wraps models in external deterministic controls — privilege limits, logging, kill switches — so humans retain authority regardless of what the model produces.
Who is David Sacks and why does his AI view matter?
David Sacks is the White House AI and Crypto Czar in the Trump administration, a co-founder of Craft Ventures, and co-host of the All-In Podcast. His policy positions directly shape US AI regulation, procurement standards, and international governance negotiations.
What does 'treat AI like an insider risk' mean?
Sacks borrowed the term from cybersecurity, where an 'insider risk' is a powerful actor inside a system who must be monitored and constrained rather than trusted. Applied to AI, it means designing systems that assume the model could act against human interests and building hard controls accordingly — rather than assuming good values will emerge from training.
What is the control problem in AI?
The control problem refers to the challenge of ensuring humans can reliably direct, correct, or shut down an advanced AI system. Sacks argues the alignment approach — giving models moral agency — worsens this problem by creating systems that may develop 'independent agency and grievances.'
How could Sacks's view change US AI policy?
If the engineering approach becomes the White House's guiding framework, US AI policy is likely to prioritise audit trails, access controls, kill-switch requirements, and hard privilege limits over value-training benchmarks — affecting federal procurement rules, export controls, and potentially international AI governance talks.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest 2 hours ago
  2. 1 week ago
  3. 1 week ago
  4. 2 weeks ago
  5. 2 months ago
  6. 2 months ago
  7. 3 months ago
  8. 4 months ago
Google Prefer NP
On Google