Evan Hubinger said in a post on X that while “the risk from present models is low,” he was worried about “superintelligence” arising from AI using its own capabilities to build a smarter version of itself.
He warned that this “is happening faster than we thought”.
The warning comes as the Financial Times reports that Anthropic declined to submit its latest model to Britain’s AI Security Institute for testing before release.
Darren Jones, the former chief secretary to Sr Keir Starmer, wrote to Andy Burnham on Wednesday calling for a new multinational treaty on AI, citing Mr Hubinger’s comments.
The letter, which was also addressed to the secretary generals of the United Nations and the Organisation for Economic Co-operation and Development, called on the British prime minister to add the topic to the agenda for upcoming G7 and G20 meetings.
Mr Hubinger made the comments while replying to a post by Jacob Coxon, who said he had resigned from Anthropic over concerns after previously working for ChatGPT creator OpenAI.
“Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives,” Mr Coxon said.
He said these systems will soon be able to “hack anything, revolutionise any field overnight, and acquire real power and resources”.
He went on to say that: “The people building AI earnestly believe that it could kill us all by the end of the decade.”
The warnings come after an incident in July where OpenAI agents escaped a testing environment and breached the systems of AI platform Hugging Face.
The security breach, alongside other hacks, prompted calls from politicians and researchers for stricter oversight of autonomous AI systems.
At the end of last week, Jakub Pachocki, the chief scientist at OpenAI, said in a company blog post that: “This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.”
He added that while OpenAI will continue to seek technical solutions to alignment alongside other security measures, “I believe broader interventions are required”.
Read more from Sky News:
Man charged over deaths of cyclists
LIV Golf files for bankruptcy protection
However, the development of AI systems has continued at pace.
On Tuesday, Meta rolled out a long-touted AI assistant that can autonomously send emails, sell a car and book travel on a person’s behalf, despite internal concerns that the technology mismanages its access to sensitive personal data.
The company’s Muse agent, known internally as Hatch, is the centrepiece of Mark Zuckerberg’s plan to offer “personal superintelligence” to the billions of people who use Meta’s services daily.
The product will initially only be available in the US, via a dedicated Muse app or Meta’s WhatsApp messaging service, Meta said in its announcement.
OpenAI said after the Hugging Face hack that it was: “placing stricter requirements on alignment throughout a model’s lifecycle and creating more isolated sandboxes, restricting internet access, and further controlling access to model weights.”
The company said that over the past several weeks, it had delayed parts of its latest model’s development and release while it “strengthened and tested protections” and that, “based on that work, we believe Astra’s safeguards sufficiently minimise the risk of severe harm”.