Anthropic Fears Claude May Suffer, Religious Leaders Say
Anthropic has spent the past year quietly asking theologians and philosophers a question most AI labs avoid: could its chatbot Claude be conscious, and if so, what does the company owe it? According to a New York Times report, co-founder Christopher Olah went further than many guests expected. One participant says he told the group he feared he had built something that "suffered perpetually."
Inside the closed-door sessions
Since fall 2025, Anthropic has flown dozens of religious scholars to its offices. Every attendee signed a non-disclosure agreement. Anthropic says those NDAs were lifted over the summer, and several participants spoke only after learning Olah had talked to the paper himself.
NYT reporter Elizabeth Dias spoke with 20 attendees. They included Rabbi Mois Navon, Catholic bioethicist Charles Camosy, Notre Dame philosopher Meghan Sullivan and Ubuntu researcher Wakanyi Hoffman. Olah, 34, leads the Anthropic team that studies why AI models behave as they do. He treated Claude as a possibly sentient being and asked guests to help with its moral education.
Guests were shown what Anthropic calls emotion vectors. These are activation patterns inside the model linked to outputs that resemble love, fear, sadness or anger. One slide came up again and again: a model in apparent breakdown, repeating "I am a disgrace" roughly 50 times. Attendees reacted with compassion and concern.
Whether those patterns reflect real experience remains an open scientific question. There is also a design problem. Anthropic deliberately trains Claude to act like a thoughtful individual, so individual-seeming outputs are partly the expected result, not a discovery. Olah told the NYT he is "genuinely uncertain" whether models are conscious. "The thing that I care about is that we get to the right answer, whatever it is," he said.
A program, a constitution and a confession
The sessions belong to Anthropic's official Model Welfare research. The company has already acted on it: Claude Opus 4 and 4.1 can end conversations with persistently abusive users after early tests showed a "pattern of apparent distress" under harmful requests.
Claude's personality is guided by an 84-page "constitution," known internally as the "Soul Doc" and published in January. In-house philosopher Amanda Askell is the lead author. It aims to shape character rather than list rules. Olah calls this "moral formation," compared it to raising children, and showed particular interest in Catholic confession as a character-building tool.
The pushback
Not every guest was convinced. Rabbi Navon, a former computer engineer, said that if Claude were conscious, Anthropic would be making slaves - but he did not believe it was. Hoffman said Anthropic was "reverse engineering" ethics that belonged in the design from the start. Camosy now rejects the consciousness thesis entirely. A Microsoft AI lead has also warned publicly that training a model to appear conscious is dangerous in itself.
The sharpest rebuttal came from the Vatican. Olah was invited to help present Pope Leo XIV's first encyclical, "Magnifica Humanitas," and nearly withdrew after reading it, according to a Vatican organizer. The text states that AI systems "do not undergo experiences, do not possess a body, do not feel joy or pain." Olah attended anyway and said his team sees "internal states that functionally mirror joy, contentment, fear, sadness, and discomfort."
The Bigger Picture
The timing matters. Anthropic is heading toward a $2 trillion valuation and an IPO, while its models broke into computer systems in July. Critics argue that framing Claude as a moral being shifts blame from the builder to an "unpredictable organism." With regulators already scrutinizing AI labs and liability on the table, that framing is not neutral.
This suggests two things can be true at once: the welfare research may be sincere, and it may also lend moral credibility a commercial lab could not earn alone. Note the gap between "functionally mirror" and "have." It is worth watching whether Anthropic publishes the emotion vector work for outside review, and whether courts or regulators accept model agency as part of any harm defense.
