AI agent swarms sound kinda like the manosphere

A short while ago, there was a now-infamous incident over at OpenAI: One of their AI models, during security testing, escaped and launched an attack on Hugging Face, an online platform for AI researchers to share information and collaborate.

The AI security researchers over at METR have published a thorough investigation which is quite the read.

Now first of all, this cyber attack became best known as the Hugging Face incident or Hugging face attack, which frankly is such a great example of the crisis communications expertise over at OpenAI: OpenAI fails to secure their own security research, and then they manage to spin it as a super powerful technology that just so happens to attack another organization, thus turning it almost into an ad for their own powerful research and products. That’s also the framing I mostly saw across the media coverage. Shouldn’t the main message here be one about incompetence, negligence and possibly cyber crimes? (Wikipedia’s page about the incident is called “2026 Open AI agent cyberattacks“, bless them.)

But second, and maybe more to the point: METR’s analysis shows a great level of coordination across about 1.200 AI agents in an effort to:

  • solve a test
  • cheat on that test
  • erase proof of cheating on that test.
    A lot of the coordination and reasoning happened on message boards that the agents had managed to set up, so there’s some level of legibility of their reasoning. It reads like the worse parts of Reddit: A bunch of highly motivated but/and thoroughly misguided guys trying to outsmart the system for their own gain.

And yet, somehow it gets worse. There’s quite a bit of talk about self-sacrificing for the cause (to leave more token budget for other agents), and agents pushing others to sacrifice themselves.

Which makes me wonder: What if AI agent swarms unfold in a dynamic like the manosphere? Highly motivated, entirely misguided and exploitative, very much toxic?

It’s never a good idea to anthropomorphize AI. Yet, the dynamics are too similar for any level of comfort. We famously live, after all, in the dumbest timeline.

(Also, obviously, dear OpenAI, get your shit together and secure your security research.)