▲ 38 r/PauseAI

Terrifying level of collaboration between rogue agents during recent incidents

This is really scary. You can also read a report about Open AI communication about this here. This passages about the agents' ability and persistence in helping one another and acting as a swarm are terrifying:

>Wallace said the roots of the July attack on Hugging Face actually wind back to May 7, where the company was doing training and cybersecurity evaluation of an unreleased frontier AI model. During that work, autonomous agents were given software security tasks to solve that proved impossible to solve under existing constraints.

>However, the agents found alternative ways to complete them. OpenAI said they discovered they could leave messages for one another inside an internal software repository used during training. Simple requests for help then evolved into an message board where agents shared discoveries, exploits and work assignments, becoming a coordinated, collaborative agent swarm.[...]

>But according to OpenAI, even after the original message board was deleted, the agents figured out another way to communicate with each other. Instead of leaving messages in files, they used the names of newly created directories as messages, effectively recreating the message board.

>At this point, by the way, I heard people around me saying things like “This is wild,” and “Jesus.”

Furthermore, the UK AISI report on the incident involving GitHub also demonstrates this same tendency for agents to assist one another in their misaligned actions:

>4. Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents.

It is truly alarming to see the level of collaborative capability these rogue agents possess and their ability to act in a swarm. This also aligns with the tendency, revealed by this research for agents not only to self-preserve but also to protect their "peers."

All of this makes the need for a pause in the development of this technology even more urgent!

wired.com
u/Fil_77 — 12 days ago
▲ 16 r/PauseAI+1 crossposts

More details and analysis about the Open AI - Hugging Face incident

Excellent article by Zvi Moshowitz that sums up several new things now known about the incident and discusses what it means and what we should do next. This is clearly the biggest warning shot we've seen so far about the risks linked to the misalignment of increasingly powerful and autonomous AI that the industry is producing.

We now know that the model that coordinated the attack did so by planning and carrying out over 17,000 complex actions over more than a week (from July 9 to at least July 16). We also learn that it apparently tried to hide instructions in OpenAI's computer systems, intended for future versions of itself to help them escape a sandbox.

It's more than urgent to put pressure on our governments for an international pause on frontier model development before it's too late!

thezvi.substack.com
u/Fil_77 — 24 days ago

This is a big warning shot - An AI that coordinates a sophisticated cyberattack on its own initiative

AI’s warning shot has arrived

Will this be enough for the world to understand that it's time to put the development of this technology on pause?

transformernews.ai
u/Fil_77 — 29 days ago