When a company adds a board member, the announcement is usually a formality. Paul Christiano's appointment came with a personal essay about the chance of losing control of AI. On the Moonshots podcast, host Peter Diamandis, founder of XPRIZE and Singularity University, said the statement read "more like a warning shot." The episode was recorded on September 16 and published on September 19. In it, Diamandis and his co-hosts argued about how seriously to take the warning and what one safety researcher inside OpenAI's governance can actually change.
Three different roles
OpenAI announced on September 9 that Christiano would join the board of the OpenAI Foundation and its Safety and Security Committee, which is chaired by Zico Kolter. The appointment involves three distinct roles:
- A seat on the Foundation board. According to the announcement, the Foundation is a separate charitable organization. It controls OpenAI Group PBC, the company that runs OpenAI's business, and holds a substantial equity stake in it. A PBC, or public benefit corporation, is a for-profit company that is legally required to consider a stated public mission alongside profit.
- Membership of the Foundation's Safety and Security Committee.
- A non-voting observer role on the PBC board. He can observe the company board but cannot vote on its decisions.
On the podcast, Diamandis described the appointment more simply, as a seat on the nonprofit board that governs OpenAI's benefit corporation.
Christiano is well known in AI safety. Alignment research, his field, is the effort to make AI systems reliably pursue the goals their developers and users intend. OpenAI's announcement credits him with foundational work on reinforcement learning from human feedback (RLHF). In this method, people rate a model's answers and the model is trained toward the kinds of responses they prefer. Diamandis called it the technique that made ChatGPT usable. The announcement also says Christiano led alignment research at OpenAI from 2017 to 2021 and founded the Alignment Research Center. It cites his government work evaluating frontier models at CAISI, the U.S. Center for AI Standards and Innovation, and its predecessor.
The feedback loop Christiano fears
In his personal statement, Christiano sets out a specific mechanism. Once AI systems can do much of the work of AI research themselves, each improved system can help build the next one faster. That creates a feedback loop in which progress accelerates. His concern is that capabilities could advance faster than the methods for keeping systems under control. His own estimate of when this could happen ranges from months to years.
Diamandis read passages from the statement aloud on the show. They included the view that the AI industry, OpenAI included, is not currently on track to reduce the risk of a catastrophic loss of control to an acceptable level. They also included the expectation that building superintelligence without more robust alignment would mean losing control of it permanently. Diamandis also read a passage citing OpenAI's own prediction that AI might have capabilities sufficient to fully automate AI research within 18 months. According to that passage, the six months after full automation could bring more algorithmic progress than the nearly ten years since the transformer, the design behind today's large language models.
The statement explains what loss of control could look like. Christiano links it to AI agents trained to pursue rewards, which might acquire resources, undermine human oversight or hide what they are doing. He gives his own estimates of the risk: 4% over one year and 15% over three years. He says these are uncertain personal beliefs, not the output of a precise model.
He calls for stronger safety measures, transparent evidence and common standards across developers, including a willingness to slow down unilaterally when necessary. He also says that joining the board is neither an endorsement of OpenAI nor a specific criticism of its practices. He asks that AI developers be judged by behavior that outsiders can verify.
"Why aren't we hearing the heads of the labs?"
Diamandis's first reaction was alarm, then frustration aimed at the industry's leaders. He asked why Sam Altman, Dario Amodei, Elon Musk and Demis Hassabis were not publicly promising to prioritize alignment. He suggested they could put half of their capacity into it and aim to build "the most aligned AI on the planet." He also proposed a way to get there. Today's models learned from "all the vitriol" of Reddit and Facebook. Instead, he suggested, labs could build synthetic training datasets free of that material, and release future models only if they were trained on data aligned with humanity's needs.
Later in the segment, Diamandis called Christiano's statements "pretty alarming and pretty extremist." Still, he concluded that Christiano means them. He said he could not imagine a board member of what he called an arguably $1.5 trillion company saying such things "without actually fully authentically believing it." He also speculated that OpenAI's decision to delay its IPO (its stock-market listing) was "probably driven" by the nonprofit governance board. That was his guess, not something established on the show, and the IPO was discussed as a separate story.
Dave Blundin: an engineering problem, not a movie plot
Dave Blundin, founder and general partner of Link Ventures, rejected the premise. "Just from a core engineering point of view, it's just not that hard," he said. He dismissed the idea of an AI that suddenly turns evil and escapes unnoticed, which he compared to the film Age of Ultron, as "nonsensical from an engineering point of view."
He agreed with Diamandis that early models absorbed junk from Reddit posts and tweets. But he called removing it "so straightforward." You get the model to produce a nasty message, trace backward through its activations (the internal signals of the network) to the training data that triggered it, and clean that data up. He also argued that nothing requires a model to have motivations at all. Much of human evil, he said, comes from drives for power, sex and food that an AI does not share. Trained the right way, it would be "more than happy" to spend its future making people happy.
In Blundin's view, the drama is partly cultural and partly about attention. The escape scenario "is not the real threat, but it's what gets the news." The terrorist scenario, he said, is "very real" but does not make headlines.
Alex Wissner-Gross: "old Paul" should meet "new Paul"
Computer scientist Alexander Wissner-Gross, founder of Reified, said he knows Christiano and thinks highly of him. He credited Christiano with championing "prosaic alignment." This is the view that there is no single magic algorithm for making AI safe. Instead, systems can be aligned in a messy but scalable way through large amounts of data and many humans interacting with them, as RLHF does. Wissner-Gross said Christiano was "exactly correct" about that.
He granted that timelines may be short. But he argued that none of the remedies he was hearing properly balanced capabilities against safety concerns. He also suggested there is a temptation to say something provocative in order to motivate one's work. His alternative is to apply the same formula at a larger scale, a "prosaic super alignment" in which superintelligent systems police one another. "Old Paul, who was advocating for prosaic alignment, should meet new Paul," he said. He added that this is what OpenAI should probably be aiming for.
What can one board member change?
Salim Ismail, founder of Open ExO, set aside the question of how large the risk is and asked about the appointment itself. He wanted to know "what information or authority or escalation capabilities come along with this appointment." In other words: what can Christiano see, what can he decide, and whom can he alert if he finds a problem? His second question was sharper. What could Christiano do that would make OpenAI "do something differently next month that it wouldn't otherwise do?"
The hosts did not answer those questions on the show. The published record helps with part of them. Christiano sits on the Foundation board, which controls the company, and on its Safety and Security Committee. On the company's own board, he can observe but not vote. His statement also offers a standard for judging the result, one that fits Ismail's test: judge developers by the behavior outsiders can verify.