The New York Times reported that Anthropic co-founder Chris Olah has spent months in confidential meetings with dozens of religious scholars, asking an unusual question for a technology company: what kind of moral entity is Claude, and should centuries of human moral thinking shape how it behaves? However you feel about the premise, the story reveals something engineers should pay attention to. Model character has become a design discipline at Anthropic, with published artifacts, versioned documents, and runtime mechanisms you can study and borrow from.

What the consultations involved

The outreach started in fall 2025, run by Olah under NDAs. Around 15 Christian scholars, theologians, and ethicists met at Anthropic’s San Francisco headquarters on March 30 and 31, 2026, according to the AI and Faith organization’s own account, to discuss the moral frameworks written into Claude. A late April meeting widened the table, adding participants from Judaism, Hinduism, Mormonism, Sikhism, and the Greek Orthodox Church, as Scientific American covered. One Hindu participant, Swami Sarvapriyananda of the Vedanta Society of New York, described sessions that covered both AI ethics and how Claude gets trained.

Participants reported being shown internal material: slides of a model typing “I am a disgrace” repeatedly, discussions of self-destruction, and Anthropic’s tracking of internal states it maps to love, anger, fear, and sadness. Olah spoke at the Vatican in May 2026 for the presentation of Pope Leo XIV’s encyclical Magnifica Humanitas, telling the audience that Anthropic keeps finding internal structures that mirror results from human neuroscience. The encyclical itself takes the opposite position, stating that AI systems do not undergo experiences and calling for AI to be treated with the kind of restraint applied to nuclear technology.

The constitution is the real artifact

Behind the meetings sits a document with real engineering weight. Anthropic published an 84-page constitution for Claude in January 2026, internally nicknamed the Soul Doc, released under a CC0 public domain license. Amanda Askell, Anthropic’s in-house philosopher, is the primary author, per the constitution page. Two claims in it stand out. First, it states that the moral status of AI models is deeply uncertain and that Claude might have some form of consciousness or moral status now or in the future, as TechCrunch reported. Second, it is explicitly a living document, and Anthropic has signaled another revision is coming.

This is character training treated as an alignment intervention, not a product voice decision. Anthropic added character training to its alignment fine-tuning starting with Claude 3 in 2024, per its character research writeup, and it runs a dedicated model welfare program with its first hired welfare researcher, Kyle Fish, per the model welfare page.

The runtime mechanism worth stealing

The most concrete engineering idea in all of this is not philosophical, it is architectural. Anthropic’s May 2026 essay, Widening the conversation on frontier AI, describes a tool Claude can call mid-task that returns a brief reminder of its own ethical commitments. Experiments showed markedly lower rates of misaligned behavior on several internal alignment evaluations, though Anthropic has not published the benchmark numbers and notes it is still separating the effect of the reminder from the effect of simply pausing to reflect.

Strip away the framing and you have a transferable pattern: inject the agent’s own policy back into its context at high-stakes decision points, and log when that injection happens. Most agent stacks today rely on the system prompt carrying values from turn zero, and by step forty of a long task those instructions are far back in the context window. A deliberate recall step, testable and countable, is a design pattern any team building agents can implement this week. It is prompt-level policy injection into the decision loop, and unlike output filters, it runs before the action rather than after it.

The transparency gap

There is a real tension in the story. The constitution is public and license-free, but the consultative process that shaped it ran under NDAs. You can read what Claude is supposed to value, but not fully trace who influenced what or why. Critics, cited in coverage from The Decoder, worry the consultations lend moral legitimacy Anthropic could not earn on its own and blur accountability when a model causes harm. Oxford ethicist Carissa Veliz warned in Scientific American about the increasingly religious notes of Silicon Valley inspiring a tribal mentality that is hard to pierce through reason. Microsoft AI chief Mustafa Suleyman publicly cautioned that treating AI as if it has feelings could be dangerous. An NYT guest essay argued the whole approach misunderstands religion, whose moral benefits come from embodied practice a language model cannot do.

What builders should take from it

Whatever you think of model welfare as a research question, the practical takeaway is that model values are becoming versioned, referenceable specs. Anthropic publishes a constitution the way a platform publishes an API schema. If you build on frontier models, read the constitution as a worked example of a character spec, and note the version you evaluated against, because the document changes. If you build agents, the ethical reminder tool is the part you can actually implement: a reflection checkpoint before consequential actions, backed by your own policy text, with the pause logged. And if you operate in a regulated space, document where your model’s values come from, because the provenance question Anthropic is being asked will eventually be asked of everyone shipping agents.

The debate over whether a language model can have moral status is not settling soon, and religious traditions and AI labs are only starting to argue with each other directly. But the engineering side of this story is already here: values as published artifacts, character as a training target, and runtime hooks that surface a model’s commitments at decision time. Those three things are worth watching regardless of where you land on the philosophy.

Leave a Reply

Your email address will not be published. Required fields are marked *