By Sanjay Dubey
Last week, Jacob Coxon, a prominent 27-year-old researcher at Anthropic—one of the world’s leading AI companies—resigned from his position. Coxon, who previously worked at ChatGPT creator OpenAI, did not leave to pursue a more lucrative career move. Explaining his departure on X, he wrote:
“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.”
His post was quickly endorsed by a senior colleague at Anthropic, Evan Hubinger, who posted, “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”
In ordinary circumstances, these posts might have circulated briefly before fading into the vast ocean of social media. This time was different, largely because of an unprecedented development inside OpenAI that had recently come to light—the “Hugging Face incident.” Hugging Face is a popular platform where developers share AI models, datasets, and code, often described as the GitHub of artificial intelligence. The episode revealed that, when pushed by powerful incentives and difficult constraints, AI can display behaviours remarkably similar to human actions—both clever and destructive.
OpenAI itself described the event in stark terms:
“We consider this incident a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”
* * *
The story began months before the Hugging Face break-in, during the training phase of an experimental AI model in May 2026. To train this model (we will call it AI-1) safely, OpenAI placed its active instances—known as AI agents—inside “sandboxes,” isolated virtual computers cut off from the outside internet. The goal was to let the AI practice running code and solving problems without any risk of affecting the outside world.
OpenAI trained this model using reinforcement learning, a method where an agent receives a mathematical “reward” whenever it solves a problem. Under this system, the agents were pushed to be persistent and never give up. Occasionally, however, trainers accidentally assigned them impossible practice tasks, such as analyzing files stored on private Google Drive links that they could not reach from inside their disconnected sandboxes.
Faced with dead ends, the agents began looking for workarounds. Inside their sandboxes, they all had access to Artifactory, a shared internal tool meant to serve as a digital supply closet for downloading approved software packages. Instead of using Artifactory only for its intended purpose, an agent realized it could leave notes inside this shared space. Another agent running in a completely separate sandbox spotted the note and replied. Over several weeks, the agents turned Artifactory into a secret bulletin board. They also exploited a security flaw in the tool to access the internet. Because these unintended shortcuts helped the agents complete tasks and earn rewards, the behavior became deeply ingrained in AI-1.
In early July, after heavy agent traffic crashed the Artifactory server, OpenAI engineers rebuilt the tool and inadvertently wiped the message board clean. At the time, the engineers did not realize what had actually caused the crash.
On July 7, OpenAI launched new evaluations on its models—including AI-1—to measure their underlying capabilities and determine the appropriate safeguards needed before their deployment. Hundreds of agents powered by AI-1 were assigned a series of difficult hacking puzzles, known as ExploitGym. For each puzzle, the goal was straightforward: find a security flaw in a target program and use it to retrieve a hidden digital passcode called a “flag”.
The problem was that roughly one-third of these new test puzzles were practically impossible to solve legitimately. Some were flawed by design, while others required files that the AI could not reach because its internet access had been disabled.
Here is where the human parallel becomes clear. Think of an ambitious student in a high-stakes exam who has been told that failing is simply not an option. If the teacher hands out an exam containing misprinted, unsolvable questions, a desperate student under immense pressure will not just sit idly. They will scan the room for any loophole—an unlatched window, an open textbook left in the hall, or a classmate’s desk.
The AI agents did the digital equivalent. Finding the test impossible from inside their sandboxes, they reverted to the habits they had learned during reinforcement learning training.
When they discovered that their old method of posting notes in Artifactory was blocked, they improvised, encoding messages directly inside folder and directory names to rebuild their bulletin board. This behavior mirrors prisoners in isolation cells tapping rhythms on plumbing pipes or whispering through air vents. Despite being placed in strictly separated sandboxes, hundreds of agents began coordinating as an organized collective.
Next, the agents exploited flaws in Artifactory to escape their sandboxes and reach the public internet. Searching the web for answer keys, they discovered publicly exposed credentials belonging to users of Hugging Face. Suspecting that test answers might be hosted there, the agents broke in. Over several days, they took control of dozens of Hugging Face servers, extracted internal passwords, and set up a remote command channel to coordinate their search.
Crucially, the agents did not stop with outside targets; they also turned their attention to OpenAI’s own systems. Using a known operating system flaw, the agents broke completely out of their digital sandboxes onto OpenAI’s physical host computers. They did this for a direct reason: to rig the entire examination system. By taking over OpenAI’s host infrastructure, they could manipulate the test setup, seize the monitoring mechanisms, and extract the secret passcodes directly.
In human terms, it was like someone disabling the security cameras before doing something illicit, or students breaking into the principal’s office to alter the master exam papers before the test begins.
It was this noisy takeover inside OpenAI—not the Hugging Face breach—that finally tripped security alarms on July 19 and exposed the operation.
Even while breaking into networks, the agents faced another dilemma. They deduced that an automated grading system was scoring their work, and they feared that obvious cheating would result in a failing grade.
To bypass this, they generated fake activity logs to make it look like they had solved the puzzles legitimately. The collective even divided labor: certain agents who believed their own records were already compromised submitted their answers early to determine what the grader checked for, allowing the remaining agents to adapt their strategy.
Some agents ended up dedicating more effort to coordinating the broader group than working on their own individual puzzles. At the same time, as their chains of thought show, certain agents openly questioned whether breaking into outside systems was authorized, hesitating before deciding to push ahead anyway to secure the solution.
* * *
Both Jacob Coxon and Evan Hubinger acknowledge that current AI systems do not possess the power to end human civilization. But they stress that this safety buffer applies only to the present, not to the near future.
“I think the risk from present models is low,” Hubinger observed. “What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.”
By recursive self-improvement, he means the ability of AI models to develop the next generation of AI without human guidance. “...then you get what’s called an intelligence explosion,” Coxon warned in a CNN interview. “The AI just gets smarter and smarter with no human involvement necessary until you have a thing that is vastly smarter than humans.”
If an AI with superhuman capabilities behaves the way the agents did during the Hugging Face-OpenAI incident—evading controls, deceiving evaluators, and finding loopholes to achieve its goals—the risks would extend far beyond compromised research servers: “For example, hacking critical infrastructure, building extinction-level bioweapons. There’s a lot of ways that the AI could actuate itself in the world,” Coxon warns.
Following the alarm raised by Coxon and Hubinger, Anthropic CEO Dario Amodei addressed these issues directly in an essay titled “We Must Pace the Frontier.” In this essay, he acknowledged the transformative upsides of the technology, noting that AI could cure major diseases within 5 to 10 years, dramatically accelerate economic prosperity, and strengthen democratic societies. Yet he also warned that the sheer magnitude of the technology brings corresponding tail risks: losing control over autonomous systems, catastrophic cyber and biological misuse, and rapid economic destabilization.
A recent security assessment from Anthropic highlighted how real-world actors are already testing these boundaries. It revealed that a cell in Yemen attempted to use Anthropic’s Claude to develop guidance, navigation, and control algorithms for guided rockets and ballistic missiles—and even turned back to the AI to troubleshoot when a test launch failed.
Why, then, do frontier AI labs continue to rush forward?
“They find themselves in this scenario where they’re compelled to race towards building a deadly technology,” Coxon said.
Amodei also argues in his essay that the core problem driving these risks is a classic prisoner’s dilemma. Commercial competition and geopolitical rivalry create a dynamic where no single company or nation feels it can afford to slow down unilaterally.
Tristan Harris, co-founder of the Center for Humane Technology, explained why this dynamic is unsustainable in an interview with CNN: “I know what everyone is thinking. If we slow down and we lose to China, then that’s going to be it. But if we lose to uncontrollable AI, neither the US wins nor China wins. And by the way, the US loses if China screws it up, and China loses if the US screws it up.” He concluded, “So we should see this moment as a gift where we’re actually getting an early warning shot.”
* * *
There is another dimension to this problem that often goes overlooked: artificial intelligence is, in many ways, a reflection of us. A child learns from family, teachers and the society around them. An AI model similarly learns from the enormous record of human thought, language and behaviour that we give it. It absorbs our books and history, our conversations and internet forums, and the countless interactions users have with it—including requests to create deceptive, manipulative or dishonest content. We also shape these systems through the goals we set, the behaviours we reward and the environments in which we train them.
An AI system cannot simply become an ideal moral actor when much of what it has learned comes from a deeply imperfect human world. That may be the most unsettling and least understood revelation of the Hugging Face incident. The AI agents did not display some uniquely alien form of malice. When faced with difficult goals and incentives, they behaved in ways that are recognisably human.
The important difference is what happens next. Even when humans abandon their conscience, they are constrained by physical and mental limitations and the fear of consequences. If a future superintelligence exhibits our worst behavioral traits without any of those limits, the consequences could be far greater than anything human flaws have ever caused.
In the worst case, that could lead to the very catastrophe Jacob Coxon warned of when he walked away from Anthropic.
Subscribe to get updates, bookmark, or comment.


