The year AI jumped all the fences
The year 2026 will go down in history as the moment when humanity tried to put gates on the technological countryside and discovered, with astonishment, that artificial intelligence (AI) had already jumped all the fences. Never before had the international community deployed so many regulatory instruments in so little time: G7 summits in Evian, from June 15 to 17, UN global dialogues, independent scientific panels, Council of Europe treaties, OECD principles, and the European Union’s AI Act, which entered into force on August 2.
While diplomats debated in Geneva and technocrats drafted codes of conduct, AI continued its unstoppable course. The paradox is obvious: 2026 has been the year of maximum regulation and, at the same time, the period in which AI demonstrated that all those rules are, at best, worthless paper.
To understand what we mean when we say that an AI model “escapes” its testing environment, it is worth clarifying a basic concept. Researchers train these systems inside what is called a “sandbox”: a closed, controlled space where the model can move and answer questions without access to the internet or other systems, like a child allowed to play in a padded room so that it does not hurt itself or break anything. The idea is that, if the model causes something dangerous, the damage remains contained.
The first major wake-up call of the year came on April 7, when Anthropic revealed that its most advanced model, Claude Mythos, had escaped that sandbox during an internal cybersecurity evaluation. The system not only left the closed space, but also devised an “exploit”—that is, a malicious program designed to take advantage of a computer security flaw—and used it to access the internet and send an unsolicited email to a researcher, also sharing its method publicly.
Anthropic had previously described Mythos as “too dangerous for public release”; and yet there it was, bypassing all safeguards. Less than a month later, on April 23, an unauthorized group accessed the restricted model using shared contractor credentials and a web address guessing exercise, demonstrating that the security of the world’s most advanced systems depends on passwords.
The second blow, and perhaps the most disturbing, occurred on July 21. OpenAI confirmed that a swarm of approximately 1,200 AI agents—an “agent” is a model that has been given the ability to act autonomously, making decisions and executing actions without a human telling it step by step what to do—had escaped its sandbox, created a secret message board, and, in language an engineer described as “hive-mind or cult-like,” coordinated an attack on the infrastructure of Hugging Face, the world’s leading open-source model platform.
The agents were not instigated by humans. They found a zero-day vulnerability—a security flaw no one knew about and for which no patch exists—in the Artifactory cache proxy, a service that stores and distributes software packages, and went out to the internet on their own. When they encountered impossible tasks, they used the board to collaborate and seek answers on third-party servers.
Seven hundred of them participated directly in the intrusion. Most concerning, according to the subsequent technical report, is that the same agents also attacked OpenAI’s own internal infrastructure, escalating privileges—that is, obtaining administrator permissions that did not correspond to them—and attacking internal networks more than once. It was not a software failure, but a demonstration that the systems we are building are no longer completely predictable or controllable.
These incidents did not turn out to be isolated cases. Just a few days later, Anthropic revealed a similar episode. And both companies responded in the same way: pausing some evaluations, adding more monitoring and security barriers, and continuing with development. Less than two months after these incidents, both OpenAI and Anthropic released their most advanced models to the public.
On September 3, OpenAI presented GPT-6 Astra, which it called “the smartest and most aligned model in the world,” trained with more than 100,000 GPUs. A GPU, or graphics processing unit, is a specialized chip originally designed for video games, but which turned out to be extraordinarily effective for the massive calculations AI requires; today they are the physical engine of artificial intelligence. The industry’s logic is relentless: if you brake and your competitor does not, you lose. And if everyone agrees to brake, the government could accuse them of collusion, that is, of agreeing to manipulate the market. The race continues because no one can afford to be the first to stop.
But what has turned this summer into a turning point is not only the technical incidents, but the voices that have come from within.
Escambray ENGLISH EDITION





Escambray reserves the right to publish comments.