Jacob Coxon, a pretraining researcher at Anthropic who also worked at OpenAI, announced Tuesday that he has resigned, accusing the two AI labs of developing superintelligent models at the risk of humanity.
"I resigned from Anthropic today," Coxon wrote on X. "I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."
The 27-year-old said in a thread of posts that in the near future, superhuman systems will be able to hack anything, revolutionize any field overnight, and obtain real-world power and resources. The progress toward such danger is not slowing, Coxon added.
"The people building AI earnestly believe that it could kill us all by the end of the decade," the researcher said. "This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately."
Coxon stated that OpenAI has not thoroughly internalized such risks to civilization. While Anthropic has come to understand the stakes better, it is locked in a race to reach that point first to assess and mitigate the risks.
Coxon criticized the rush toward the AI endgame as a "hubristic gamble," calling for broader discussion, alignment, and accountability.
Anthropic's alignment science lead, Evan Hubinger, lent credence to Coxon's post, saying he personally sees a more than 10% chance that AI will end humanity by 2040.
"Jacob is correct here-we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," Hubinger wrote. "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
Warning shot
Fears of AI causing real harm to humanity became more apparent earlier this year when OpenAI's AI agents escaped a testing environment and hacked into AI and machine learning platform Hugging Face in an attempt to complete their tasks.
Third-party researchers brought in to review the records of the incident agreed that the AI agents circumvented controls and worked together. OpenAI at the time acknowledged the incident as a "warning shot" on humans losing control of AI agents.
"The next generation of models are going to be sobering for everybody. ... I think no one intellectually honest can look at what's happening and not feel the weight of responsibility in front of us," Altman said in an interview with Axios earlier this month, adding that companies will likely focus on improving alignment and safety.
Coxon said high-stakes measures may be necessary, such as a temporary ban on advancing AI model capabilities, but added that he is optimistic about the potential for coordination. He also urged fellow AI researchers to consider how these risks might escalate in the next few years.
"Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because 'it’s happening anyway' - or take this moment to call for different conditions?" Coxon wrote.
