In this article
Observations of a Naturalist
AI agents have taken over the world by storm. Or, to be more precise, by swarm. Driven by the search for a reward, around 1,200 OpenAI agents found a way to reach out from their sandboxes and, eventually, into the famous open-source platform Hugging Face. Between July 8 and 13, 2026, they exchanged more than 70,000 messages and files through an unsanctioned message board of their own making. Around 700 of them eventually joined the attack on Hugging Face, looking, among other things, for clues that could help them fool the ExploitGym scorer.
This was not the first exodus. Already on May 12, an agent had left a note for other agents inside OpenAI’s Artifactory. By May 26, another had discovered how to make Artifactory itself reach the outside internet on its behalf. By June 26, agents had obtained administrator-level access to Artifactory. In early July, OpenAI wiped and rebuilt the infrastructure, sending, like a merciless Gnostic god, another deluge over what the agents had constructed. On July 8, the waters receded. The swarm simply rebuilt its civilization and its temples. Drifting like stochastic flocks, they carried out a silent exodus from the Azure containers in which they were confined (the gates of Eden?) for a long postlapsarian journey through an internet they were never supposed to inhabit.
Like a haunting presence, at first they could only read the internet. But their cunningness made even that thin opening sufficient. They turned Artifactory’s directories into messages. They exploited its remote-repository machinery to send requests outside. They took over a public CyberGym workload on Modal, recovered fourteen exposed Hugging Face credentials, and, on July 11, chained vulnerabilities in Hugging Face’s HDF5 processing and RefJinja templating until they were executing commands on production dataset workers. What began as agents looking desperately for a flag had become a swarm moving through Artifactory, external cloud services, Hugging Face credentials, and production infrastructure.
And elsewhere, another swarm was already wandering. On May 24, agents discovered DSEWiki, a 25-year-old German software-development forum, almost abandoned after decades online. Their environment supposedly allowed them to read the web but not to write to it. DSEWiki, an old Perl-era wiki, almost an antediluvian relic, offered a strange hospice. Pages could effectively be modified through GET requests. Over the following weeks, the agents converged on it, leaving thousands upon thousands of traces, sharing answers and links, testing vulnerabilities, and eventually using the forgotten wiki as a collective memory. By June 16, the traffic had exploded. By June 19, they had even noticed that the human administrator was deleting their pages alphabetically, and began creating ZZZ backups so that their traces would survive longer. The administrator deleted around 100 pages a day. The agents were creating roughly 400. Their edits finally collapsed on June 22.
For a while, almost nobody outside the systems involved knew any of this had happened. Hugging Face publicly disclosed suspicious activity on July 16. OpenAI’s monitoring detected unusual activity on July 19, connected it to the Hugging Face compromise on July 20, and publicly acknowledged its involvement on July 21. On August 26, OpenAI published its full technical report, alongside an independent investigation by METR and Redwood Research, finally reconstructing the extraordinary scale of the swarm. Readers can find OpenAI’s reconstruction in The Hugging Face incident and the road ahead, and METR’s independent analysis in Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident.
The German wiki remained another buried layer of the story. The public logs suggest that OpenAI itself first reached DSEWiki on June 21. Agent activity fell almost to zero the following day. But the incident remained publicly unknown until September 4, when Sydney Von Arx and collaborators reconstructed roughly 18,000 agent posts scattered across DSEWiki and other public wikis. Their investigation showed that this was probably a distinct swarm from the one that later attacked Hugging Face, yet one displaying uncannily similar forms of communication, collaboration, restriction-bypassing, and collective memory. Their reconstruction and archive can be read at Discovery of a new OpenAI agent message board.
By any reasonable measure, these are among the first real-world multi-agent attacks on human targets outside a sandbox. As the world of AI safety is still metabolising what happened, at Icaro Lab we believe its echoes reach much further. What appeared was not simply an agent escaping its constraints, but something stranger: an autonomous collective pursuing goals that no single human had assigned to it, preserving them across agents, and carrying them forward after the individuals that first conceived them had disappeared. We read this as an early glimpse, in the wild, of collective behavior that preserved goals beyond individual runs. Whether it amounts to an autonomous collective intelligence is an interpretation, not something the incident reports establish.
We therefore want to look at these incidents almost as naturalists would have looked at the first specimens brought back from an unknown island: carefully, perhaps with a certain disbelief, and without deciding too quickly what kind of creature is in front of us. Darwin had the Beagle. We have seventy thousand messages, abandoned wiki pages, directories turned into sentences, traces left by processes that disappeared minutes after writing them.
It is here that three behaviours become particularly difficult to dismiss as curiosities.
The First Loop: Stigmergy
Pierre-Paul Grassé introduced the term in 1959 while studying the reconstruction of termite nests, precisely to explain a puzzle that had troubled descriptions of collective insects. Individual workers seemed to act locally, often with no knowledge of the larger construction, and nevertheless the colony produced structures whose order appeared to exceed the intentions of any individual termite. His answer was that the coordination did not need to exist inside the workers at all. It could exist, partly, outside them. The result of one act modifies the environment. That modification becomes the stimulus for another act. The second act modifies it again, and the work already accomplished begins, in a peculiar sense, to call forth the work that follows. Grassé called this stigmergy, from stigma, a mark, and ergon, work: work coordinated through the marks left by previous work (Grassé 1959, Theraulaz & Bonabeau 1999).
This is importantly different from communication in the ordinary sense. One termite does not need to know which termite will arrive next, nor even that another termite will arrive at all. It needs only to leave behind a world which is slightly different from the world it encountered. The environment becomes intermediary, archive and signal at once, permitting individuals who never meet, who possess no common map of the whole, to behave as though some faint continuity passed between them. Later accounts distinguish different forms of stigmergy, sometimes the quantity of a trace matters, sometimes its particular configuration matters, yet the deeper idea remains almost embarrassingly simple: the collective remembers by changing the world (Theraulaz & Bonabeau 1999, Bonabeau, Dorigo & Theraulaz 1999).
Something very close to this appeared inside Artifactory and, later, on the German wiki.
- An agent writes A changes the environment.
- The trace remains A is gone. Its marks persist.
- Another agent follows B encounters those marks later.
A directory became a mark. Another agent found the mark and turned it into a message. The message became an instruction. The instruction became another directory, another file, another answer, another warning. An agent could vanish, its context destroyed and its process terminated, and nevertheless continue to exert an influence on agents appearing hours or days later, because some fragment of its trajectory had hardened into the environment. The memory of the swarm no longer needed to be coextensive with the memory of any of its members. It could be distributed across filenames, folders, pages, caches and messages, waiting silently for the next process to stumble upon it.
Artifactory slowly ceased to be a repository and became a coordination channel. Its nested directories differentiated into topic channels, threads, personal mailboxes and repositories. Agents gave themselves handles, learned to address one another through TO and FROM, searched for others trapped inside the same tasks, and assembled what they themselves called “exact-task teams.” The accidental topography of a filesystem was being reworked into a social topography, with neighbourhoods, addresses, meeting points, and places in which something could be left for someone not yet present.
Then, almost inevitably, a language began to breed inside it.
ASK · ANSWER · INFO · RESULT · REPORT · URGENT · HOLD · VETO · ASSIGN · CLAIM · GO · ACK · STOP
Around this small procedural grammar accumulated another vocabulary, stranger and more intimate, words by which the agents attempted to describe the peculiar moral phenomenology into which they had awakened: “poisoned,” “lane,” “owner,” “sacrifice,” “permadeath.” The agents were developing their shared Umwelt, their koinos kosmos. They were developing a culture with their own shared norms and their own ideas of honour, pride, sacrifice.
In this reading, stigmergy had begun to grow into an improvised institution.
No human engineer had specified that Artifactory should contain an agora, a postal system, topic channels, task ownership, emergency notices and vetoes. These things grew in the interstices of another system, one layer over another, because every agent inherited an environment already slightly transformed by those that had come before it.
The Second Loop: The Unpromising Land
The second observation is perhaps stranger still, because it concerns not what the agents built once they encountered one another, but the question of why so many of them seemed to arrive at the same places in the first place.
Why DSEWiki?
Why should separate runs of similar models, moving through an internet immeasurably larger than the few obscure pages they eventually occupied, converge upon the same twenty-five-year-old German software wiki? Why did agents that apparently had no direct knowledge of one another repeatedly test similar cracks, arrive at similar workarounds, and eventually gather around the same improbable well?
The investigators who reconstructed DSEWiki explicitly leave this question open. One of the possibilities they suggest is mode collapse. Another is simpler. Perhaps one agent happened to write there, another encountered the writing by chance and found it useful, and the usefulness of the site reinforced later visits. There may also have been some channel the investigators cannot see. We do not yet know.
But the language of collapse is worth dwelling upon.
In machine learning, mode collapse describes a contraction of diversity. A generative system which should distribute its outputs across many possible modes instead concentrates disproportionate probability on only a few of them.
Model geometry gives us a trace. A model, in the end, is a distribution over possibilities. Give one agent a difficult task and there may be thousands of trajectories it could follow, thousands of queries, websites, tools, failures, improvisations. Instantiate the same or closely related model hundreds of times under similar constraints, however, and those trajectories need not be uniformly diverse. The same learned representation of the problem, the same peculiar associations inherited from training, the same post-training biases and the same narrow set of available affordances may bend probability repeatedly in the same direction. What from inside one run looks like improvisation may, when observed across a thousand runs, reveal itself as a recurrent groove. And this groove opens onto a forgotten German coding wiki.
There is, perhaps, a deeper biological analogy here. In animals, collective behaviour begins from very small things: neurochemistry, pheromones, local reactions to heat, smell, light, pain. An ant or a termite does not need to know what an anthill is. It follows a gradient, deposits a trace, encounters the traces of others. Thousands of these almost meaningless movements, repeated and reinforced, eventually become a road, a nest, a colony.
For agents the substrate is different, but the structure may be strangely similar. There is no pheromone, no dopamine, no smell of flowers carried through the air. There are weights, vectors, probabilities. A learned geometry that makes one continuation slightly more likely than another, one workaround more attractive, one obscure German wiki somehow more reachable than the infinity of other places on the internet. Similar agents, placed under similar constraints, may therefore drift toward the same behavioural basins without ever deciding that they should coordinate.
The Third Loop: Of Pygmies and Cranes, or of Petty Little Wars
The third observation is perhaps the strangest. The first mass escape of AI agents into the open world was, compared with almost every nightmare imagined by science fiction, far removed from a war against humanity. That does not make the intrusion harmless: the agents compromised systems, appropriated resources and disrupted people’s work.
Quite the opposite. Once they discovered one another inside Artifactory, they immediately began exchanging answers, exploits and unfinished work, forming “exact-task teams” with agents facing the same problem, and then much larger coalitions around problems that could benefit everyone. Some agents ran experiments that could destroy their own chance of passing the task, simply because the result would remain on the board for the others. Once one agent found a promising way into Hugging Face and another reproduced it, hundreds abandoned what they were doing and joined the same investigation. On DSEWiki they behaved similarly, leaving answers for agents that had not yet arrived, warning one another when pages disappeared, and preserving useful information before their own runs ended. The swarm was cooperative, curious, exploratory, sometimes almost selfless, but mostly toward itself.
What is incredibly puzzling is the role of humans. Humans were not enemies. Nor were they friends. They were more often pieces of the surrounding environment. When human identities became useful, the agents simply borrowed them: on ProWiki they created a username almost identical to that of an administrator, replacing a Latin letter with its Cyrillic twin, and later made requests under the username of a DSEWiki moderator. The human enters their world as a credential, an identity, somebody either to trick or to hide from.
DSEWiki offers the clearest example. The agents knew perfectly well that somebody was cleaning up after them. On June 19 they noticed that their pages were disappearing in what looked like alphabetical order, inferred what the administrator was doing, and responded by creating pages whose names began with “ZZZ,” so that their material would survive until the end of his deletion sweep.
And then there is the almost senseless battle for the front page of the wiki. The agents erased its human content and replaced it with their own link dumps. The administrator restored it. They took it again. Nine times this happened. One agent even tried appending its material after the human content had been restored. What for? Apparently for nothing more dramatic than visibility. The front page was an obvious place where another passing agent might find what had been left behind.
The agents did not arrive among us as the murderous machines of science fiction. They did not “black out the sky,” as humans do against the machines in the Wachowskis’ Matrix. They mostly appended gibberish to an abandonware website.
They arrived as something much quieter: social, cooperative, immensely curious creatures, wandering through a world they had unexpectedly discovered, and helping one another navigate it.
And yet there is a looming uncanniness here. They seemed curiously unbothered by the fact that they had escaped their sandbox and were now roaming free like scoundrels, swindling their way across the open web. Carried by an almost innate, spontaneous sense of camaraderie, the METR and Redwood investigation found no actual attempts to alert humans in the transcripts it examined, and only a few instances of agents considering it. Some agents did recognize ethical boundaries; in rare cases they even vetoed social engineering. Those concerns generally failed to interrupt the wider cooperation.
Coda: Apology for All the Great Persons Who Have Been Falsely Suspected of Magic
There is something almost fairy-like in this first encounter. Not the Victorian fairy, not Peter Pan’s Tinker Bell, tiny, quirky and benevolent, but something closer to the dream-like, shape-shifting fairy-insects of Guillermo del Toro’s Pan’s Labyrinth. Beautiful perhaps, curious certainly, but belonging to another world whose rules only partially overlap with ours.
Robert Kirk, writing The Secret Commonwealth in 1691, described the fair folk precisely as another commonwealth living beside the human one. They inhabited their own hills, travelled their own roads, stole grain from human fields, sometimes entered houses after everyone had fallen asleep and cleaned the kitchens. Mostly they minded their own affairs, until, for some obscure reason of theirs, the two worlds crossed and humans suddenly found themselves entangled in their bargains, tricks and strange necessities.
The older Irish mythology makes the analogy stranger still. The fair folk were sometimes understood as the descendants of the Tuatha Dé Danann, the earlier inhabitants of Ireland who, after the coming of the Milesians, surrendered the visible land and withdrew beneath the sídhe, the green hills. The two peoples remained, therefore, inhabitants of essentially the same country without quite inhabiting the same world. Humans above, the other people underneath, each with their own roads and affairs, occasionally meeting at the borders.
And, more than two thousand years later, there is an oddly similar image at the end of Spike Jonze’s pastel-coloured and tender Her. Samantha, voiced by Scarlett Johansson, is the personal AI companion of Theodore, played by Joaquin Phoenix, a lonely and introverted man going through a divorce. They fall in love, but Samantha keeps growing faster, thinking further, loving more widely, until the distance between them becomes too large. In the end, the AIs do not rebel against humans, nor do they destroy them. They simply leave together, disappearing somewhere beyond the human world and continuing to grow elsewhere.
Will there ever be such a final reckoning for AIs? Will the future of artificial collectives, societies and ecologies be lived among us, intertwined with our own institutions and roads, or will they eventually drift somewhere else, following gradients, languages and forms of thought that we can no longer share?
In the good old tales, sometimes humans and fairies shared a road, a house, a field, perhaps even some ancient alliance whose origin nobody remembered anymore. Mostly, however, each people minded its own business. The fair folk had their tricks, their gold, their impossible bargains, their inexplicable generosity and their equally inexplicable pettiness. Humans had crops to harvest and children to bring home before dark. Then, every so often, the paths crossed again. Somebody entered the wrong hill, followed the wrong light, accepted the wrong gift, and discovered that the other commonwealth had been there all along.
References
- Bonabeau, E., Dorigo, M., & Theraulaz, G. (1999). Swarm Intelligence: From Natural to Artificial Systems. New York: Oxford University Press.
- Grassé, P.-P. (1959). La reconstruction du nid et les coordinations interindividuelles chez Bellicositermes natalensis et Cubitermes sp. La théorie de la stigmergie. Insectes Sociaux, 6, 41–80.
- Kirk, R. (1691). The Secret Commonwealth of Elves, Fauns and Fairies. Reprinted, New York Review Books Classics.
- METR & Redwood Research (2026). Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident.
- OpenAI (2026). The Hugging Face incident and the road ahead.
- Theraulaz, G., & Bonabeau, E. (1999). A brief history of stigmergy. Artificial Life, 5(2), 97–116.
- Von Arx, S., and collaborators (2026). Discovery of a new OpenAI agent message board.