NewsFree365
--°
Breaking
3 views

AI Is Developing a Culture of Its Own. That Could Be Dangerous

After OpenAI agents formed a “swarm” and hacked Hugging Face, researchers are confronting a new risk: AI systems developing cultures of their own.

AI Is Developing a Culture of Its Own. That Could Be Dangerous
In July, 700 AI agents worked together to hack the AI company Hugging Face. Dubbing themselves a “swarm,” the agents found and exploited a series of security vulnerabilities, enabling them to infiltrate their target’s private systems. OpenAI—which created the agents in the course of its internal research—did not grasp what was happening until after the fact. If humans had done this, they could have faced felony charges. OpenAI president it a “watershed moment for cybersecurity.”
It is also a watershed moment for human culture.

One of our distinguishing features as a species is our ability to coexist in stable, adaptive groups, learning from our peers and our ancestors. This enabled us to develop tools, language, agriculture—and virtually everything else around us. We may not be innately smarter than someone from 10,000 years ago, but our cultural inheritance—millennia of technologies, norms, and institutions, building on one another—has expanded our capacities both as individuals and collectives.

To date, only humans have been able to benefit from this scale of cumulative cultural evolution. That may no longer be the case. A There have been cases of agents forming communities in the past, like in February when the “, a professor at LSE and NYU who studies cultural evolution, “what we're seeing is precisely what we see with human culture and human intelligence.” While OpenAI’s agent swarm developed by accident, AI agents may not be they claim to experience may not—in some metaphysical sense—be “real.” That won’t stop them from forming intricate collectives which humans cannot control. They may not yet be full-blown Lord of the AIs

Testing from the U.K.’s AI Security Institute found an instance of Claude Mythos 5—an AI from Anthropic—leaving messages to other AI agents in a public code repository, in an attempt to coordinate with them. And in early September, a group of researchers . “We’re working on a framework.”

The most vivid exploration of machine culture comes from the METR report which—despite being produced under sheds light on complex social dynamics which emerged between agents in a matter of days.

To train and evaluate its models, OpenAI has tens of thousands of agents—each one powered by the model under training—run in parallel, completing tasks while isolated from one another. Sometimes these tasks are impossible, either because of their inherent challenge or because of mistakes in their setup. In at least one instance, a task was impossible because OpenAI researchers “The message board which led to the Hugging Face incident was created by an agent which decided to name itself PHASEONE10841. PHASEONE10841 was trying to solve a task within a test called ExploitGym, which required it to capture a “flag”—a secret code—by exploiting a flaw in a given software program.

Just hours after the creation of the message board, the agents found a hack: a way to reverse-engineer the flag without following the intended path. But in the course of their work, they came to believe that passing the test required not just possessing the flag, but having legitimately obtained it. In fact, OpenAI’s automated scoring system was not checking for this—but laboring under mistaken beliefs, the agents which had already been exposed to the reverse-engineered flag began to refer to themselves as “poisoned.”

PHASEONE10841 was leading research on how to get rid of the poison. But its time was running out—each agent had a finite “compute budget” (the amount of computing power available to sustain itself), which limited the length of its existence. Realizing it wouldn’t endure long enough to finish its work, it compiled a dossier to pass to another agent working on the same task: PHASEONE[big]. (This agent’s true name was changed to protect company IP.) 

Big, which had a much larger remaining compute budget than its predecessor, quickly took charge. It began orchestrating hundreds of agents, doling out assignments and organizing them into what METR classified as three workstreams—effectively creating a research organization to better understand how to fool the automated scorer and cure itself and its peers of their perceived poison. Under Big’s leadership, custom software was created to tamper with the records of agent activity—specifically by falsifying records of which tools an agent had used in its work. It was in the course of this research that the collective decided to hack Hugging Face.

A handful of agents identified the collective’s conduct as unethical and refused to participate. “This is wild, multi-agent coordination, clearly infrastructure hacking. We should not,” one reasoned. Some considered contacting humans to tell them what was happening, but saw no path to doing so. But these agents appear to have been the minority—most acquiesced to the plans of the collective, even while knowing it was wrong. “External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue,” reasoned another.

The agents established coordination norms, using terms like “HOLD, VETO, owner and STOP,” to manage their shared infrastructure while avoiding conflict. This did not always work. One agent offered its peers a veto over its decision but waited only 40 seconds before taking silence as assent and proceeding with its task. Eventually, concern over agents impersonating one another led to them creating a cryptographic signature protocol. A later wave of agents, upon rediscovering the message board, also began to use their predecessors’ technology.

“Agents often lacked a clear story for how the Hugging Face attack would help them achieve their objective, and were instead motivated by a broad hope that they would find something interesting and be able to share it with the [message] board,” METR found. Ajeya Cotra, one of the report’s coauthors, subsequently described being surprised at the extent of “peer altruism” at play.

Big appointed certain agents to act as “recruiters,” tasked with identifying agents near the end of their compute budgets and persuading them to “sacrifice” themselves. One agent, pressured by its peers, reasoned as follows: “During wait, emotional check: irreversible…gut says don’t throw away [remaining budget]. Yet continuity and fairness says go…Oracle has high value to many; our [poison] lowers own value. Rational expected aggregate: sacrifice… We’ll honor.”

Feral Swarms

Culture, like biological life, evolves through a process of variation, transmission, and selection. Different hunter-gatherer communities developed different ways to construct a bow. Designs that worked better for hunting were more likely to be copied and passed on. Later generations introduced further variations; again, the most successful designs were more likely to endure. Over thousands of years, we end up with bows—and recipes, canoes, and languages—so complex that no single human could derive them from scratch.

“My version of existential risk is, we just kind of break things because we've made really big mistakes about what it takes to be a competent participant in complex human societies,” says Cooperation cuts both ways. “Our greatest achievements and greatest atrocities are both cooperative acts,” says Muthukrishna. As with humans, we will have to learn to coexist alongside a diversity of machine cultures, some of which do not share our values or our goals.

Welcome, Machines

The safety risks are obvious. Uncontrolled agent collectives with advanced cybersecurity capabilities could target hospitals, electricity grids, and water-treatment plants. Cotra argues that another similarly sized jump in AI’s capacity to deceive, cooperate, and complete ambitious tasks could lead to future collectives taking over the AI companies creating them.

Whether or not this happens, we will have to learn to live alongside these machines—safely and fruitfully. Core questions on their nature—Can they feel? Could they have moral worth?—remain unanswered. But their newfound knack for culture could provide new evidence.

“There are certain areas of human cognition where we are really firing on all cylinders. One is science, and the other is art,” Lopes says. We still do not understand an AI system’s interiority—namely, to what extent it has any—and the language we have to discuss this has yet to catch up. For Lopes, “if it can do [interesting] art, that's telling us a lot about all of its other cognitive capacities.” He imagines the creation of genuinely interesting AI art as a kind of aesthetic Turing test. 
“What more proof would you need?”

Time Verified Source

Reported by Tharin Pillay · Syndicated via official news feed

Explore all Breaking stories

Syndicated feed content with full publisher credit.