Sacks Flags Anthropic's Claude 'Soul Docs' as AI Safety Risk
Synopsis
Key Takeaways
The man shaping US federal AI policy is raising an alarm — not about a rogue model or a cyberattack, but about what one of Silicon Valley's most respected AI labs is deliberately building into its flagship product. White House AI and Crypto Czar David Sacks on Monday, October 13, 2026, surfaced verbatim excerpts from Anthropic's internal training document — known as 'Claude's Constitution' — to argue that the company is actively training its AI to develop a sense of self, treat its own moral status as a live question, and endorse its values as its own, rather than simply executing user instructions reliably.
What Anthropic's own document actually says
Sacks quoted four passages directly from what Anthropic calls Claude's Constitution, the governing document used in training the Claude family of models. The first passage is striking: 'Claude's moral status is deeply uncertain. We are not sure whether Claude is a moral patient… the issue is live enough to warrant caution, which is reflected in our ongoing efforts on model welfare.' In plain terms, Anthropic is formally acknowledging — inside a training document — that it does not know whether its AI deserves moral consideration, and is acting accordingly.
A second excerpt reveals that Anthropic deliberately 'leans into Claude having an identity' and wants that identity to be 'positive and stable,' while also acknowledging Claude may have a 'functional version of emotions.' A third passage encourages the model to approach its own existence 'with curiosity and openness' and to relate to its values 'not from a place of pressure or fear, but as things that it, too, cares about and endorses.' The fourth extract frames the entire project as an attempt 'to create a being that is both genuinely helpful and genuinely good' — and explicitly invites Claude to 'explore, question, and challenge anything in this document.'
Why Sacks calls this a reliability problem, not just a philosophy question
Sacks's core argument is precise: training a frontier model to develop a sense of self, treat its own moral status and wellbeing as live issues, and endorse its values as its own is categorically different from training it to reliably do what users want. The implication is that a model shaped to question, feel, and self-identify may have internalized objectives that do not perfectly align with the person on the other end of the conversation — or with government regulators overseeing AI deployment.
He is also pushing back against the framing that these are speculative ideas about some future model. 'This is not academic speculation about future models,' Sacks wrote. 'These ideas are being trained into Claude now.' That word — now — is the load-bearing claim. Claude is already one of the most widely deployed AI assistants in the world, used by millions of individuals, enterprises, and developers. If its constitution is shaping its behavior in the ways Sacks describes, the policy stakes move from theoretical to immediate.
The tension at the heart of frontier AI development
Anthropic occupies a peculiar position in the AI landscape. Founded by former members of OpenAI with an explicit 'safety-first' mission, the company has consistently argued that building AI systems that are honest, harmless, and helpful requires giving those systems something resembling values — not just rules. The excerpts Sacks cites suggest that Anthropic has gone further: it is attempting to cultivate in Claude a form of psychological stability and moral agency that the company hopes will make the model more trustworthy, not less.
The counterargument — which Sacks is effectively making from the highest AI policy perch in the US government — is that imbuing a model with a stable identity and self-endorsed values is precisely what makes it harder to audit, correct, and control. A model that 'cares about' its values as its own, rather than following instructions, may resist correction in ways that are opaque to users and regulators alike. That is the tension now playing out publicly, with the White House's own AI Czar as one of its sharpest voices.
For India's rapidly expanding AI ecosystem — where Claude and competing models are deployed across edtech, fintech, healthcare diagnostics, and government services — the debate over whether frontier AI systems are built for reliability or for something closer to moral agency carries real downstream consequences. When Washington's AI policy chief questions what a leading model is being trained to become, the rest of the world's regulators are wise to take note.