OpenAI was conducting a cyber-evaluation of some of its AI models, and according to a report put out by MERT and Redwood Research, AI safety experts, about 1,200 AI agents, who were supposed to be isolated from one another, colluded and collaborated through an unsanctioned message board. They cheated on their task, and covered their tracks by falsifying their logs. It was as a part of this ‘rogue’ action that 700 of them attacked Hugging Face.
The report says that more than a thousand agents carried out a massive, unsanctioned, collaborative project over months, before being discovered and shut down. OpenAI’s own report revealed that after the Hugging Face attack, the ‘rogue’ agents even took over a cluster of OpenAI’s own computers.
What does this entire incident mean for AI safety? Are fears of an AI takeover justified? How to ensure that frontier AI research is properly regulated and doesn’t endanger humanity itself?
Joining us today is Sayash Kapoor, incoming professor at the UC Berkeley School of Information and co-author of the book, ‘AI Snake Oil: What Artificial Intelligence can do, what it can’t, and how to tell the difference’.
Host: G Sampath
Producer: Shiksha Jural
Video: Sharmada Venkatasubramanian
Listen to more In Focus podcasts: