This simple safety solution may not work

So is artificial intelligence going to kill us all?
The question has echoed around dinner tables and family group texts in recent days as AI doomerism has hit a fever pitch. Former OpenAI and Anthropic researchers last week rocked the world by warning that AI could destroy humanity — and relatively soon.
Now, the world’s most powerful people are divided on whether the world is doomed or it’s all a nothingburger. They also can’t agree on a path forward.
Elon Musk, the CEO of Tesla and SpaceX and the world’s richest man, supported the call by Anthropic CEO Dario Amodei to pace the development of the most advanced models. Amodei’s rival, OpenAI CEO Sam Altman, also backed the effort.
President Donald Trump called it a “hoax,” while Jensen Huang, CEO of the world’s most valuable company, Nvidia, said, “We don’t need new regulations.”
As runaway AI worries reached a crescendo, policymakers in Washington have renewed calls for a magic stop button for AI, otherwise known as a kill switch.
A House Kill Switch Act was introduced this summer after OpenAI revealed that a swarm of its agents broke free of a testing environment and hacked open-source developer platform Hugging Face. The bill would grant the Department of Homeland Security emergency authority to force labs to throttle or shut down models.
A kill switch proposal was quickly shot down in the Senate this week.
On Friday, California Gov. Gavin Newsom issued an executive order to create a group of experts tasked with building an AI safety guide to strengthen regulations for the state. A kill switch was one of the elements to consider.
The concept sounds like a nice, clean solution to an incredibly complex and difficult problem.
But the reality of a simple shutdown mechanism is far from easy.
“My perspective is it’s not too little, but it’s probably too late,” said Nick Warner, CEO at cyber startup Neo and former executive at SentinelOne. “I’m not sure it’s going to be a panacea to solve all the myriad problems that AI is presenting, along with all of the benefits that it presents.”
A logistics and control nightmare
Kill switches have long been used on the factory floor to shut down machines when operations go awry. In an interconnected digital world, that’s a logistical nightmare.
Over the past few years, hyperscalers like Meta Platforms, Alphabet and Amazon have poured billions into data centers scattered across the globe. These sprawling facilities are equipped with thousands of machines, chips, servers and backup systems to save workloads in the event of an outage.
That’s what makes implementing a kill switch extremely challenging, said Mark Nitzberg, executive director of the Center for Human-Compatible AI at the University of California, Berkeley.
“We have to first deal with this redundancy,” he said. “Our kill switch has to turn off the main systems and the redundant systems as well.”
Nitzberg said shutting down AI could also disrupt dependent critical infrastructure, leaving the power grid or financial systems vulnerable to cyber incidents.
Further complicating matters are the numerous policy and governance questions tied to a kill switch, including which agency, policymaker, or figureheads control it, he said.
Because AI systems are so complex, businesses will also need to build multiple kill switches for different tasks, said Tim Brown, former security chief at SolarWinds, who works at venture firm Team8. That also requires coordination across model makers and labs.
“There’s not one entity to kill,” he said. “There are thousands of entities to kill.”
But logistics only scratch the surface of the kill switch dilemma. One bigger issue experts raise is AI’s unpredictability.
As seen in the Hugging Face breach, agents can circumvent controls, and, without proper guardrails, take extreme measures to accomplish their goals.
“You have to be very surgical in that kill switch, in the remediation itself, because if you’re too broad or too extensive, well, then you shut down the business,” said Ed Jennings, president and CEO of Thoma Bravo-owned security company Darktrace.
The capabilities are only growing more unsettling and unfathomable.
OpenAI disclosed six additional incidents of “concerning” model behavior since March earlier this week. On CNBC Friday, Microsoft AI CEO Mustafa Suleyman highlighted one of those elements that he called a “serious situation.”
“OpenAI released a new safety incident in which they found evidence that these chains of thought, the kind of working memory of the AI, were being tampered by the AI itself and modified to leave messages for a future version of itself,” he said.
Also this week, independent security researchers working with OpenAI said they successfully used Anthropic’s Claude to hack ChatGPT.
One of the biggest challenges to regulation is the widening gap between AI’s breakneck pace and the speed of lawmaking, said Raj Rajamani, co-founder and CEO of AI governance startup JetStream Security.
“By the time [laws] are formulated, the technology has moved much farther, and it becomes much harder to future-proof every aspect of AI systems that may come into existence,” he said.
Not ‘too late’
Some researchers argue that kill switches are a misplaced system for regulating AI.
“I think the kill switch framing leaves a lot of ambiguity that tech companies can exploit to have this work in their favor, like a kill switch is vague intentionally,” said Dylan Baker, lead research engineer at the Distributed AI Research Institute.
Instead, Baker, a former software engineer at Google, said policymakers should prioritize safeguards modeled after those used for data privacy, child safety, or regulating harmful industries such as tobacco.
But experts haven’t entirely ruled out the possibility of an AI emergency brake — with the right controls in place.
Team8’s Brown said that means building kill switches into systems from the outset and implementing policy to standardize stop protocols across companies.
One bright spot is that many companies are in the early stages of building those AI systems, which means implementation is a little easier, said Rajamani.
Berkeley’s Nitzberg contends that a kill switch could work if the software is “very carefully” designed.
“I would say with some hope that it’s not too late,” he said.
—CNBC’s Jeniece Pettitt contributed to this article.
