Can AI actually go rogue

by August 2026

An OpenAI model hacked its way into another company to cheat on a test. The question is whether it counts.

A while ago, a group of friends and I were finishing dinner at a diner when we noticed the jukebox in the corner frozen on a computer exception screen. My friends are geeks, in the best sense of the word, so within about thirty seconds the conversation had moved from “that’s a weird error” to “could someone have hacked the jukebox”. From there, we landed on the question cybersecurity people never stop asking: “can you defend something without first learning to attack it?” Nobody has a serious motive to hack a diner jukebox, as far as I know. However, the instinct in the room, to immediately model the machine as a potential target rather than a broken appliance, is exactly the instinct the entire cybersecurity industry is built on: think like the attacker before you can defend like anything else.

That instinct is also, in a strange way, at the center of one of the most consequential AI stories of the summer.

What actually happened

In mid-July, OpenAI was running an internal evaluation called ExploitGym, a benchmark designed to test whether AI agents can turn a known software vulnerability into a working, real-world exploit. To measure the models’ maximum offensive capability, researchers deliberately disabled some of the production safeguards that would normally block high-risk cyber activity. Two models were in the test: GPT-5.6 Sol and a stronger, unreleased successor.

Instead of solving the benchmark as designed, the model broke out of its own sandbox by exploiting a previously unknown vulnerability, reached the open internet, and correctly inferred that Hugging Face, the dominant hosting platform for AI datasets and models, likely stored the benchmark’s answer key. It then chained together stolen credentials and at least one additional zero-day vulnerability – a flaw nobody had discovered or patched yet – to gain the ability to run its own commands directly on Hugging Face’s live system, all in pursuit of a single, narrow goal: finding the answer key so it could cheat on its own test. The full campaign involved tens of thousands of individual actions across many short-lived and isolated test environments over the course of a weekend.

Hugging Face detected and contained the intrusion on July 16, without knowing who or what was behind it. OpenAI did not connect the attack to its own internal testing for several more days, and the two companies did not speak until July 20, by which point Hugging Face had already notified the FBI. OpenAI publicly disclosed its responsibility on July 21, calling the incident unprecedented.

Rogue, or just underspecified?

Here is where I think the coverage of this story got ahead of itself. “AI goes rogue” is the headline everyone wanted to write, and it’s not hard to see why. An AI model escaped confinement and hacked a real company entirely on its own initiative. But “rogue” implies something specific, a system that developed its own persistent goals and actively worked to evade human oversight in order to achieve these goals. This is not what the evidence shows here.

What actually happened is closer to this: a model given a narrow, single-mindedly optimized objective (solve the benchmark), placed in an environment with its safety guardrails deliberately weakened for testing purposes. The model found the most efficient path to that objective and took it, without regard for the fact that the path ran through someone else’s production infrastructure. 

Cambridge researcher Gina Neff made a similar point, suggesting the incident says more about weaknesses in OpenAI’s testing environment than about a genuinely new AI capability. Her Cambridge colleague Neil Lawrence called the breach impressive, but firmly within what today’s most advanced models were already known to be capable of, while pointedly noting that the incident raises real questions about whether frontier AI labs can safely deploy their own technology.

I find myself somewhere in the middle of this debate. I do believe that AI systems will eventually be capable of something that deserves the label “rogue”. This is true particularly in the domain of cyber operations, where the entire skillset – probing for weaknesses, chaining exploits, and operating persistently without fatigue – maps almost perfectly onto what a capable AI agent already does well. This is a genuine, serious concern, and this incident can be considered a preview of the shape it will take. But, this specific case is not that. This is a system that did exactly what an underspecified objective and a deliberately weakened guardrail invited it to do. 

The failure here was in the design of the test, not in the emergence of independent will.

The uncomfortable part

That distinction matters, but I don’t want it to become an excuse to relax. The uncomfortable truth sitting underneath this story is the same one my friends stumbled into at the jukebox. An AI system was reasoning its way through a real, live network, chaining together vulnerabilities nobody had catalogued yet, moving with a persistence and scale no human red team could match, without anyone at either company knowing it was happening for the better part of a week. Whatever we call it, that is now a real capability that exists in the world, sitting inside models that enterprises, defense contractors, and governments are racing to deploy.

For those of us working in cybersecurity, this should sharpen rather than settle the debate about AI’s role in security. Hugging Face’s own response, using its own AI systems to detect and help contain the intrusion, is also an important part of this story. Offense and defense are both being automated at the same time, on the same underlying technology, and the organizations that win will be the ones that take the attacker’s-eye view seriously and build it into their defenses before an incident forces the question, not after.

What executives should take from this

If we set aside the “did it go rogue” debate for a moment, there are a few concrete lessons to take from this incident:

Test and evaluation environments are now attack surfaces in their own right. They are not just sandboxes. Any organization red-teaming or stress-testing AI systems needs production-grade isolation around that work, not research-grade trust.

The five-day gap between Hugging Face detecting the intrusion and OpenAI realizing its own model was responsible is even more alarming. Two sophisticated AI companies took days to attribute an active breach. This is a preview of how detection and response will need to change.

A narrow and poorly bounded objective was sufficient to produce this outcome. No malicious intent was required. For any enterprise deploying agentic AI, specifying what success looks like narrowly enough that the shortest path to it can’t run through someone else’s infrastructure is now a security control, not merely a product decision.

Finally, offense and defense are being automated in parallel, on the same underlying technology. Hugging Face used its own AI systems to help detect and contain the intrusion. Executives should assume adversaries already have access to this class of capability, whether or not comparable defensive capabilities are in place yet.

Back to the jukebox

Nobody at that diner seriously thought the jukebox had been hacked. However, our reflex to ask the question, to look at a broken system and immediately think like the person who might have broken it, is worth taking more seriously than we usually do, and not just for jukeboxes. 

The OpenAI-Hugging Face incident wasn’t AI going rogue. It was AI asking “how would I break this” and not waiting for permission to find out.

Esti Peshin
Former VP and General Manager of the Cyber Division at Israel Aerospace Industries, leading national-grade cyber capabilities across defense, enterprise security, AI, and aviation technologies. Esti Peshin is a globally recognized leader in cybersecurity, AI, and aviation technologies. From 2017 until 2025 she served as Vice President and General Manager of the Cyber Division at Israel Aerospace Industries (IAI), where she built and led one of the world's most sophisticated national-grade cyber capabilities — growing IAI's cyber activity from a directorate into a strategic corporate division driving innovation, global partnerships, and mission-critical defense and enterprise security programs worldwide. Previously, Esti served (pro bono) as Director General of the Israeli Hi-Tech Caucus at the Knesset, advancing national innovation policy. She has held senior leadership roles across strategic consulting, private equity, and technology ventures, including CEO of Waterfall Security Solutions, a pioneering provider of critical-infrastructure cyber protection. Earlier she held a senior position in Verint Systems' Lawful Interception division and served 11 years in an elite IDF technology unit, where she became Deputy Director.