Business & Finance

OpenAI's new safety hire says losing control of AI would be 'catastrophic' and that 'most people could die'


OpenAI’s new safety hire didn’t mince words about what’s at stake in the age of AI superintelligence.

“If we build superintelligence without more robust alignment I expect we will permanently lose control of it. If that happens then most people could die,” Paul Christiano, an AI safety researcher, said on Wednesday after OpenAI announced he was joining its board and Safety and Security Committee.

In a lengthy statement about his new role, Christiano addressed the threat AI poses and warned that humans are at imminent risk of losing control of AI systems.

“Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term,” Christiano wrote. “I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level.”

Christiano previously led alignment research at OpenAI and worked for the federal government as the head of safety at the Center for AI Standards and Innovation within the National Institute of Standards and Technology.

On Wednesday, he said his joining of OpenAI was “not an endorsement or criticism of OpenAI’s safety practices in particular” but that he hopes all frontier labs improve their safety oversight. He also said reducing the risks posed by AI would require coordination worldwide.

AI researchers have consistently warned about the risks posed by superintelligence and the singularity, a term used to refer to the point at which AI systems surpass human intelligence, improve themselves, and advance at a rate that humans can no longer predict or control. OpenAI CEO Sam Altman said in July that the singularity had already arrived.

The risk of losing control of AI

Christiano said he thinks the risk of losing control of AI is so great right now due to AI’s increasing abilities in AI research, which he said could lead to a “rapid intelligence explosion.” He also said the way AI is trained with reinforcement learning means the models are being taught to “get as much reward as they can.”

“It has long seemed theoretically possible that this could motivate AI agents to undermine human control, seek power and resources, and cover up their tracks in pursuit of misaligned goals correlated with reward,” he said. “Public evidence from recent incidents suggests that this is not just a theoretical possibility.”

Christiano said that frontier companies still have the opportunity to address the situation, including by improving coordination, slowing development as needed, adopting share safety standards, and transparently sharing information about risks and mitigation efforts.

His statement came less than a day after another AI researcher shared a dire warning about AI. Jacob Coxon, an Anthropic researcher who previously worked at OpenAI, said Tuesday night that he had resigned from his role and that neither company was acting responsibly.

“They are racing straight to self-improving superintelligence and gambling with our lives,” he said on X.

OpenAI has also lost several safety researchers over the years, with some casting doubt on the company’s commitment to developing AI safely.

Please Subscribe. it’s Free!

Your Name *
Email Address *