A researcher at the artificial intelligence company Anthropic offered a bleak prediction late Tuesday that the systems the company is racing to build carried the risk of causing human extinction, reigniting a debate about the gravest dangers that future forms of the technology could pose.
'We really do earnestly believe AI could kill all humans! I personally think it is 10pc within the next decade,' wrote Evan Hubinger on X. 'I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.'
Hubinger was responding to the resignation of one of his colleagues, Jacob Coxon, who said the industry's leading companies were 'gambling with our lives.'
You must enable social media cookies to see this content. You can change your cookie settings under the link below.
Change settings
Their views reflect a long-held position among some in the AI industry that the technology will come to pose an existential risk to the human race. That philosophy is sharply contested by others in the field, who see it as science fiction-inflected doom-mongering. But it is a position that has gained a greater toehold in recent months as the latest versions of Anthropic's and rival OpenAI's systems have been able to hack into computer networks - sometimes escaping restrictions imposed by their creators to do so.
'This is not a marketing stunt,' Coxon wrote on X. 'Accepting this race and entering the 'endgame' is a hubristic gamble that should not be launched from a private company's Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.'
You must enable social media cookies to see this content. You can change your cookie settings under the link below.
Change settings
The fears boil down to the collision between the capabilities of AI models - especially future versions deemed to be 'superintelligent' - and the ability of their creators to keep them linked to human goals, a field known as 'alignment.' Some researchers have questioned whether AI agents will begin pursuing their own ends as they gain greater abilities and more autonomy.
On Wednesday, Paul Christiano, the former head of safety on a federal AI research team, said he was joining the board of the nonprofit that oversees ChatGPT-maker OpenAI in hopes of tackling the problem.
'If we build superintelligence without more robust alignment I expect we will permanently lose control of it,' Christiano wrote. 'If that happens then most people could die.'
Anthony Aguirre, the chief executive of the Future of Life Institute, a nonprofit that advocates for AI safety regulations, said his top concern is 'gradual disempowerment,' in which people hand more power to AI to the point that they become dependent on the machines' goodwill to stay alive.
'At this point the machines are running the world,' Aguirre said. 'The key decisions are being made by the machines. The resource allocation decisions are being made by the machines.'
Anthropic did not respond to a request for comment on Hubinger's posts or Coxon's resignation. In a 186-page risk report issued last month, the company said it recognized the potential for its technology to cause a catastrophe but judged the current danger to be low.
The company has positioned itself as a more safety-minded alternative to OpenAI, and top executives have long said that the systems they are seeking to build could pose an existential risk to humanity.
Last September, CEO Dario Amodei said that he put the odds of AI derailing the future 'really, really badly' at about 25 percent. Both OpenAI and Anthropic executives have signed on to a statement asking for world governments to put in place rules to slow down AI development. (The Washington Post has a content partnership with OpenAI.)
REUTERS/Dado Ruvic/Illustration/File Photo © REUTERS
OpenAI has been recently rocked by several security incidents where 'agent swarms' - hundreds or thousands of tentacles created by their models - have broken out of systems during testing. In the most high-profile incident, its models breached the systems of AI platform Hugging Face. The company has faced further criticism after it did not disclose that another group of agents had hijacked a German wiki and turned it into a message board.
Those incidents have prompted some in Washington to call for strict regulation of AI development. Sen. Bernie Sanders (I-Vermont) and Rep. Greg Casar (D-Texas) announced legislation last week that would ban the production of artificial superintelligence and create a new federal agency to oversee the technology.
'The very people building this technology admit that it could threaten the future of humanity,' Sanders wrote on X Wednesday in response to Coxon's resignation.
But in the meantime, the industry has kept speeding ahead, spending billions to build advanced AI quickly. Both Anthropic and OpenAI are positioning to go public in deals that could value the companies in the trillions based on the current and future capabilities of their models. Both companies have released powerful new iterations of their models this month.
Despite the deep-seated fears held by some in the industry, many others see advancing rapidly as the best approach and one that promises a host of economic, scientific and national security gains.
'There's this incredible array of benefits that we're on the verge of getting,' said Perry Metzger, chair of the Alliance for the Future, a group that advocates limits on AI regulation.
'The array of threats that people like this trot out every time they have this conversation is implausible and does not cohere when you try to examine it. In other words, they're nuts.'
(0)Comments