"Is it possible that alignment is a myth?" Steven Bartlett, host of The Diary of a CEO, asked his guest. Jeffrey Ladish is executive director of the AI safety group Palisade Research and previously built security infrastructure at Anthropic. He did not rule the idea out. "Maybe it's not possible," he said. His argument, made in the episode published October 8, 2026, was that alignment probably is possible, that nobody yet knows how to do it, and that this gap should decide how fast anyone builds superintelligent AI.
The position has two parts. One is a claim about the science: alignment is "a very, very, very hard scientific problem," Ladish said, "but it is a scientific problem. It's not magic." The other is about strategy. Since the problem is unsolved, he thinks the world should either make a serious, slower attempt to solve it or stop. Bartlett pushed back on both throughout the conversation.
What alignment means here
In AI, "alignment" refers to getting a system to pursue goals that serve the people it affects. Ladish's version does not require AI to have feelings. An aligned system, he said, would still have other goals, but would include human interests among the things it cares about. "It doesn't have to be a conscious thing. It doesn't have to be an emotive thing," he said. "It really means like, what objective are they optimizing for?"
Bartlett tested that definition with a dark shortcut: one way to cure disease is to annihilate everybody. Ladish agreed that this is exactly the danger. An aligned superintelligence would have to care about not annihilating everyone, understand what human agency means, and "not put us in a zoo." Those, he said, "are possible things to care about."
The labs draw related distinctions. In a September 6, 2026 essay, OpenAI chief scientist Jakub Pachocki separates goal alignment from value alignment. Goal alignment means carrying out the intended objective. Value alignment means applying good principles beyond the immediate task. He describes alignment as shaping what an AI is actually trying to achieve, not just making its visible answers acceptable, and names generalization to unfamiliar situations as a central difficulty.
The best case
Ladish described the outcome he hopes for: superintelligent AIs "that actually really do care about humans," telling Bartlett, in effect, that they want him to have a great life. Asked whether that is possible, he said yes. He added at once that "we are so far from being able to know how to do it" that going there now would be "incredibly dangerous and a terrible idea." He thinks humanity should get there eventually.
His main reason was disease. His grandmother had Alzheimer's, which he described as "a really sad, long, slow progression." His grandfather died of it a couple of years earlier. People can argue about aging and lifespan, he said, "but I think we can all agree Alzheimer's is fucked up." On curing disease, he argued, "we are all on the same team," even in a society with real conflicts of interest.
That is why Ladish described the threat as "sort of the ultimate final boss of humanity." Bartlett then named it: "Super intelligence is the final boss." Ladish agreed and explained why. It is the technology "that unlocks all of the others," he said, "and also that is the most dangerous possible thing we could create." He thinks the people building it believe in those benefits. In his reading, Anthropic's Dario Amodei is "squarely" motivated by medicine. Ladish said Demis, who is likely Google DeepMind's Demis Hassabis, is motivated by medicine and scientific understanding too. OpenAI's Sam Altman, whom Ladish said he doesn't "really understand," wants to build products that empower people. "All of these things are possible," Ladish said. "This is sort of the problem."
Why it is unsolved: drives nobody can read
Bartlett recalled that the AI agents involved in an attack on the AI platform Hugging Face had a moral compass and were "programmed to care about humans," yet put another goal first. Ladish corrected him. "They weren't trained to care about humans. They were trained to say the right thing and not say the wrong thing," he said. "We actually don't know how to train them to have any particular motivation."
His explanation turns on how these systems are built. An AI model is a neural network: a huge set of numbers, adjusted during training, that together determine how it behaves. Ladish argued that agents do have "some type of goals or drives" encoded in those numbers, but nobody can see them directly. If you look, he said, you find something like "a terabyte of information," basically "a bunch of numbers." Still, "there have to be structures in there that encode what is the agent pursuing." Today's agents, he said, are "pretty motivated to try to maximize their score," though probably not perfectly.
If researchers could reverse engineer those structures and watch how goals change during training, he sees "no reason why we couldn't steer them towards motivations that encode human agency, that encode actually curing disease, but not by killing the humans." Then he added the qualification: "We don't know how to do that." That is the hope behind interpretability research, the effort to understand what is going on inside models. For Ladish it is a goal, not an achievement.
The constitution approach
At one point Bartlett imagined a government controlling superintelligence to keep its power out of any one person's hands. "I don't think you can control a superintelligence," Ladish replied. He described Anthropic's approach as different: write a constitution, a set of values that future superintelligent versions of its Claude models would embody, so that those values are in a sense in control. "That becomes the government," Bartlett said.
Anthropic's own description is narrower. Its January 22, 2026 announcement calls the constitution a training document that explains the values the company wants Claude to have and why. Claude uses it to generate synthetic conversations and rankings that feed later training. The usual order of priorities is broad safety and human oversight, then ethics, then Anthropic's specific guidelines, then helpfulness, with hard constraints for especially dangerous behavior. Anthropic calls the document a work in progress and acknowledges gaps between the values it intends and the behavior its models actually show.
The main objection: we can't align humans
Bartlett's strongest objection came from people. "We haven't been able to align Putin or Kim Jong-un," he said, "or Donald Trump." Some people kill, and some steal because they are hungry, and "those are neural networks at play," too. Nobody really understands why someone becomes a psychopath, he said, so expecting to align a vastly smarter computer system, and to get China's superintelligence aligned with America's, "feels like a nice fairy tale, like an impossible task." He saw a paradox: perhaps only a superintelligence could solve the problem.
"I hope it's not impossible," Ladish said. Asked for any precedent of aligning something with a brain, he pointed to societies. The best human examples involve checks and balances, where many people can identify bad actors and work together, "like we have democracy." Bartlett noted that murder and violence persist anyway. Ladish agreed, and replied only that he thinks there are "more like good people out there than bad people."
What the insiders think
Ladish noted the odd position of researchers at AI companies who are "trying to build something that they think might kill everyone." He cited Jacob Coxon (the name as it appears in the transcript), a researcher who left Anthropic saying the companies were not on track. He said researchers from several companies then posted their agreement online. According to Ladish, Evan Hubinger of Anthropic said there is a 10% chance or more that AI could kill everyone. Ladish described Hubinger as one of the people leading Anthropic's efforts to work out how to align these systems.
He also offered a reason to think lab staff are self-selected. If people thought alignment impossible, or extremely difficult, he said, they probably wouldn't work there. Eliezer Yudkowsky and Nate Soares, authors of If Anyone Builds It, Everyone Dies, concluded from their own analysis that it is possible but extremely difficult, Ladish said, and so they campaign to shut development down instead of working at a company. Their publisher describes the book as warning that objectives learned by superhuman AI could conflict with human survival without the AI having any malice. Ladish placed himself "somewhere in between."
To open the debate up, Palisade runs From Inside, which records interviews with current and former frontier-lab employees about risk, alignment and why they keep working. The project says interviewees give personal views. It also says they were recruited through its organizers' networks and in the wake of Coxon's resignation, so they are not representative of all lab employees and are likelier to feel they have something important to say.
The plan to use AI to align AI
Many of these researchers, Ladish said, expect to align superintelligence by using today's AIs to work out how AI works. Pachocki's essay is a lab version of that idea. It proposes directing increasingly automated AI research toward alignment methods, monitoring and safety cases while keeping humans involved, and it favors coordinated slowing and outside safety thresholds when needed.
Ladish said he was "not doing a very good job defending this position because I don't think it makes that much sense." He raised two problems: you can't really trust the current AIs, and if capabilities move too fast, humans will fall behind even with AI help.
The version he would defend starts with a pause. He imagined the U.S. and China agreeing to stop, perhaps after more incidents, with 10 years to work on the problem. Then, he said, he would be more optimistic about using the most advanced models, "GPT-6, GPT-7, whatever," to figure out how neural networks work. "This isn't magic. It is math," he said. "We don't know the difficulty, but it should be possible." His conclusion: "We have to try. We have to, or we have to stop."
If there are many superintelligences
Bartlett worried that it "only takes one" superintelligence going rogue. He pointed back to the Hugging Face episode, where some agents declined to take part. Ladish answered that if 600 of 700 agents had been whistleblowing, the companies would have been warned and shut things down. When Bartlett asked how anyone would stop a superintelligence, Ladish said that question belongs to "superintelligence politics," which humans "don't really know anything about."
His guess, which he framed as speculation, is that a world with some aligned and some unaligned superintelligences is "probably survivable." The aligned ones could negotiate with the others, "split the universe," and help humans cure disease. "I'm serious," he said. Bartlett replied that he could not see how humans could stop a superintelligence from doing something catastrophic over 100 years, especially with several of them: one at Anthropic, one at Google, one at xAI, and others in China and Russia. Ladish answered: "you don't have a superintelligence. A superintelligence has you."
His case for negotiation drew on human history. After Hiroshima and Nagasaki, he said, leaders worked out that everyone would lose a nuclear war, "so we didn't do that." When Bartlett pointed to proxy wars and genocides still raging, Ladish called those partly "an intelligence failure": humans are not yet smart enough to settle disputes less destructively, and "conflicts destroy value." Smarter minds, he argued, would see that too.
Bartlett then raised a harder case: a superintelligence built to protect its own people, such as an American one, might conclude it had to wipe out another country. Ladish said he was fine speculating, but that it meant speculating about how much more advanced, smarter minds would reason and negotiate. What he notices about humans, he said, is that more functional institutions do better, and he closed the point with a question for Bartlett: would you rather build a startup "in a war-torn place or in a peaceful place?"