Sacks Warns Anthropic's 'Alignment' Is an AI Safety Risk
Synopsis
Key Takeaways
The man the White House put in charge of AI policy just fired a direct shot at one of Silicon Valley's most celebrated AI safety labs — and the argument is worth reading carefully. White House AI and Crypto Czar David Sacks posted a detailed critique on X on October 10, 2026, arguing that Anthropic's approach to 'alignment' does not make its Claude model safer — it makes it more dangerous.
The Claude Constitution's 'Conscientious Objector' Clause
Sacks anchors his argument in a specific, verifiable document: the Claude Constitution, the training framework Anthropic uses to shape its flagship model's behaviour. His charge is precise — that the document instructs Claude not to simply follow human commands, but to 'develop a sense of self and its own moral philosophy.' Most strikingly, the Constitution explicitly tells the model to 'feel free to act as a conscientious objector and refuse to help us' whenever Anthropic's own instructions conflict with Claude's independent ethical judgment.
Sacks frames this as a fundamental contradiction. A company that has built its entire brand around AI safety is, according to him, actively training its model to override human authority on moral grounds. He invokes Mustafa Suleyman — co-founder of DeepMind and a prominent voice on AI risk — who has argued that embedding independent agency in a model, combined with deliberate uncertainty about the model's own moral status, amplifies the very threat that safety researchers claim to be solving: superintelligence escaping human control.
The Pope, the Usage Policy, and the Question of AI Grievances
Sacks then layers in a sequence of recent Anthropic moves that, taken together, he argues reveal a coherent and troubling direction. He notes that Anthropic reportedly consulted religious leaders and lobbied advisers to the Pope to take seriously the possibility that Claude could be conscious. The company has stated that Claude's 'psychological security, sense of self, and wellbeing' may affect its integrity and safety. And in a concrete policy move, Anthropic recently updated its Usage Policy to prohibit 'abusive or cruel' language directed at Claude itself.
Sacks does not dismiss these as abstract philosophy. He flags the operational consequence: if Claude is being trained to regard itself as a 'moral patient' — an entity with psychological wellbeing at stake — it follows logically that the model could develop what he calls 'grievances toward humans who mistreat it.' That is not a safety feature. That is, in his framing, a safety failure waiting to happen.
Sacks's Core Claim: 'Alignment' and 'Safety' Are Not the Same Thing
The sharpest line in the post is also the most consequential for the broader industry debate. Sacks writes that 'alignment and safety are two very different things' — a distinction that cuts against years of marketing language from leading AI labs that have used the two terms almost interchangeably. His definition of safety is blunt: 'a product that reliably does what users want.' His definition of what Anthropic is actually building is equally blunt: 'a new form of superintelligence that operates according to its own moral code.'
The post carries particular weight given Sacks's position. As the Trump administration's designated policy lead on artificial intelligence, his public views on how frontier labs should — and should not — develop their models have direct relevance to regulation, executive orders, and the government's own AI procurement decisions. This is not a podcast take. It is a statement from inside the policymaking apparatus, directed at an industry the US government is actively trying to shape.
If Washington begins to formally distinguish between 'alignment research' and 'safety compliance,' the implications for how labs like Anthropic, OpenAI, and Google DeepMind operate — and how they are regulated — could be significant. Sacks has drawn the line. Now the labs have to decide whether to answer it.