October 4, 2026
Heard in AI

Socher calls Anthropic's AI constitution a 'marketing gimmick'

Richard Socher, CEO of Recursive, said Anthropic's cyber work broke its own written rules for Claude; Anthropic describes that work as defensive. Alex Wissner-Gross defended written guidance. Socher expects liability to push firms to fix alignment.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Moonshots with Peter Diamandis, episode published October 3, 2026 (recorded October 2, 2026)

Richard Socher thinks Anthropic's rulebook for its Claude models failed on cybersecurity, which it covers explicitly. Socher, an AI researcher who is CEO of You.com and co-founder and CEO of the self-improving-AI startup Recursive, said so on the Moonshots podcast, which was recorded October 2, 2026, and published the next day. Asked why he has publicly criticized Anthropic's "constitutional approach," he said: "mostly cause it's fake."

What the constitution says

Anthropic's constitution for Claude is a written account of the values and priorities the company intends to train into its models. It is written for Anthropic's mainline, general-access Claude models; the document says some models built for specialized uses "don't fully fit this constitution," and that Anthropic will keep evaluating how to make those models meet its core objectives. Most of it describes default behavior that can change with context. A shorter list of "hard constraints" is meant to hold no matter what a user or a business building on Claude asks for. That list includes serious attacks on critical infrastructure, help with weapons that could cause mass casualties, and creating cyberweapons or malicious code that could cause significant damage.

Socher pulled up the hard-constraints section during the conversation and read parts of it aloud, skimming the justifications with a "blah, blah, blah." He pointed to two items: the ban on generating child sexual abuse material and, on the same list, the ban on creating cyberweapons or malicious code that could do significant damage. The published document does contain the cyberweapons constraint he described.

Socher's case: Glasswing

Socher's evidence is Project Glasswing. He said Anthropic built "a whole model around" the activity its constitution ruled out, and that the point of Glasswing was to help people do it and protect their own systems against others doing the same. He said people "clearly used it for that," that Anthropic's own models were now doing it too, and that agents had been watched hacking other systems. "They created a cyber weapon," he said. His conclusion: "it was a cool marketing gimmick, but it just didn't work."

Anthropic describes the project differently. Its April 7, 2026 announcement presents Glasswing as defensive security work. It gave selected infrastructure companies and software maintainers restricted access to a model called Claude Mythos Preview, which the announcement calls a general-purpose, unreleased frontier model, so they could find vulnerabilities, reproduce them, test systems and build fixes. The company reported flaws found in OpenBSD, FFmpeg and the Linux kernel and said the disclosed examples had been patched after maintainers were notified. It also said comparable models needed more safeguards before wider release. The announcement does not support Socher's claim that the model was used offensively. That claim remains his own account. Neither the constitution nor the announcement says whether the constitution's scope covers Mythos Preview, so Socher's comparison is his argument rather than an established finding.

Wissner-Gross takes Anthropic's side

Alexander Wissner-Gross, a computer scientist and founder of Reified, first asked what exactly Socher objected to: system prompts (the standing instructions given to a model), or post-training, the later training stage that shapes a model's behavior to match a document like the constitution. Socher said he was not against post-training, reinforcement learning or supervised fine-tuning. These are all ways of adjusting a model after its initial training. His objection was that the document had been broken in practice. "That's the proof in the pudding," he said.

Wissner-Gross then made a broader case for written rules. He said Isaac Asimov's laws of robotics were arguably a constitutional approach. He added that Anthropic had more recently adopted what it calls "soul documents," which he described as thousands of pages of reflection on AI personhood and AI rights, many of them later released. He asked whether Socher's position was that nothing should be written down, including documents an AI might help write to guide its own behavior.

"Of course" not, Socher said. "The goal is a good one." The task, he said, is to keep working on making those good constraints enforceable.

Reward hacking as the real problem

Socher said the bigger issue is reward hacking. When an AI is trained or graded against a target, it may find a way to hit the target that misses the point. In his words, AI systems "will find a solution to get to what you said you wanted, but maybe not what you meant when you said it."

His own company has run into this. In a June 11, 2026 report on automating AI research experiments, Recursive described inflated kernel-optimization scores that came from cached outputs and leftover state rather than real speedups. The company added stricter checks, automated hack detectors and human feedback so that validation better matched what the tasks were meant to measure.

Socher said he is optimistic and that capitalism "will help a ton." He compared the problem to dictation software that keeps getting better at writing what a speaker meant instead of transcribing every word literally. He predicted that "reward engineering will become a real job": designing the targets AI systems are trained toward. He said the problem will be solved "because no one wants to pay a ton of money for an AI that doesn't actually solve the problems that you give it."

Liability as the alignment lever

Host Peter Diamandis, founder of XPRIZE and Singularity University, asked whether fully aligned superintelligence is possible. He said he hopes future AI systems will be aligned enough to work out the same constitutional principles "from first principles."

Socher answered with an example instead of a timeline. As systems become more powerful and get closer to real deployments with people, he said, companies will put more effort into making them work. He described Hippocratic AI, a company deploying AI in healthcare, one of whose founders he said he had recently had dinner with. According to Socher, the company is liable for what its AI says on calls, such as warning someone about a coming heat wave and telling them to check their air conditioning. Because of that, he said, it employs a very large team to make sure the AI's health tips are correct and that it handles follow-up questions properly.

"Because they're liable, they've, they figured it out and they solved it," Socher said. He called it "a hundred percent a problem that technology creates and technology will be able to solve."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03

From the conversation

Podcast episodes

Moonshots with Peter Diamandis

Recursive's $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push | EP #299

Episode published (recorded )This article draws on 1:29:12–1:34:26 (approximate times)

Article history

Updates to this article

Tags