Traditional Safety Measures are ‘Unraveling’ as AI Advances, UN Panel Warns

Traditional Safety Measures are ‘Unraveling’ as AI Advances, UN Panel Warns

The Summer of Rogue AI is almost at an end, but humanity’s understanding of how to control the technology—and prevent it from wreaking havoc in the future—is still very much in its infancy.

That’s the key message from the United Nations’ International Independent Scientific Panel on Artificial Intelligence’s first thematic brief, published this morning as world leaders gather in New York for the UN General Assembly.

The July Hugging Face hack showed that current frontier AI agents, even when handed a seemingly benign task, can establish their own dangerous sub-goals, organize into hierarchies, and try to cover their tracks to hide misaligned behavior from human researchers. Since future agents are likely to be even more capable and make decisions in ways that are even more difficult for humans to monitor, any safeguards the tech industry might put in place today will soon become obsolete, according to the 40-person Panel. “In simple terms,” it wrote in a press release accompanying its new report, “the traditional model of safeguarding is unravelling.”

Yoshua Bengio, the Panel’s co-chair and one of the so-called “Godfathers of AI,” has previously said that the best approach to ensuring the safety of future AI systems is for the tech industry to adopt the “precautionary principle,” a policy observed in many other industries which holds that all possible measures should be taken to prevent public harm, even when science has yet to prove that a product is actually dangerous. Pharmaceutical companies, for example, must undergo a rigorous clinical trial process and receive approval from the U.S. Food and Drug Administration before a new medication can be prescribed or sold over the counter. In a similar vein, two experts suggested a framework last week for how the U.S. and China could work together to prevent an AI-triggered nuclear exchange, modeled upon Cold War-era human command and control safeguards. 

The clinical trial process and analogous safety testing procedures in other sectors could provide a blueprint for the AI industry as it develops ever-more powerful systems, according to Qinghua Lu, an AI safety researcher and Panel member. “We are not starting from zero,” she said in a statement. “Aviation, medicine and cybersecurity learned to manage high-risk systems through incident reporting, independent scrutiny, and layered safeguards. But those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor.” The Panel also called for legally protected channels for AI company whistleblowers, the use of “AI-based monitoring” to detect agent misbehavior that humans might miss (though this introduces its own challenges), and “emergency intervention” measures—in other words, a quick and reliable way for researchers to shut a system down or block its access to certain tools if it goes off the rails.

Silicon Valley erupted in debate last week about the possibility that misaligned AI could one day (perhaps in the next decade) wipe out all of humanity. Doomsday scenarios have tended to be vague, but they generally converge on the idea that future AI agents could covertly take control of critical systems out of human hands, all the while assuring us that everything’s fine. The recent warnings have also sparked new public debate about AI safety, even though many of the most powerful Silicon Valley executives have openly said for years that the technology could lead to human extinction. 

According to the UN Panel, the debates over “p(doom)” and what form an AI-triggered apocalypse might take are missing the more important point. “Although the probability of loss of control events remains uncertain and the best response is still under debate,” it writes in its report, “a clear conclusion emerges: given the severity of these events, risk management requires far greater attention and resources.”

Brian Heater Avatar

Leave a Reply

Discover more from AZ Shopping

Subscribe now to keep reading and get access to the full archive.

Continue reading