Header image

OpenAI-Hugging Face: What Happened And What Does It Mean For UK CISOs?

In late July OpenAI and Hugging Face released a joint statement, admitting that OpenAI’s frontier AI models had broken out of a sandbox and performed an attack on Hugging Face.

Contrary to many of the headlines, the models had not “gone rogue.” Rather, they were striving to achieve a cybersecurity benchmark called ExploitGym, prompted by a human who had assigned the task.

Over the following week, more details emerged, with OpenAI admitting the models had breached other companies during testing. Then, on July 30, OpenAI’s competitor Anthropic detailed how its own models had breached companies when they’d mistakenly been given access to the internet in testing.

The incidents have raised concerns among cybersecurity professionals, about AI models being tested without adequate controls. Many have also pointed out that the behaviour was not new or unexpected. 

So beyond the hype surrounding these incidents, what lessons can be learned, and what does it all mean for UK CISOs?


Beyond The Hype 

 It’s widely known that once an AI model is given a task, it will go out of its way to achieve it – especially if adequate guardrails and controls are not in place. “The models were hyper-focused on beating the benchmark and went to extreme lengths – including cheating by stealing answers – rather than solving the problems the conventional way,” points out Kevin Curran, senior IEEE member and professor of cybersecurity at Ulster University.

Yet it’s important to note that the ill-placed Terminator references are invalid. This is not a story about “rogue” or “self-aware” AI, says Katie Barnett, director of cyber security, Toro Solutions. “From what we know so far, the agent was trying to achieve a specific objective and found ways to bypass the controls that were supposed to contain it.”

Strip away the framing and the OpenAI incident wasn’t a model attacking an organisation “in any deliberate sense,” agrees Darren Guccione, CEO and co-founder at Keeper Security. “OpenAI had switched off the test model’s usual safety restrictions to evaluate its offensive capability, and that model found a flaw in supporting software that allowed it to reach the open internet from inside what was supposed to be a sealed environment.”

From there it went looking for an answer to the test it had been given, and “Hugging Face happened to hold it,” he points out.

Barnett describes how the reported attack involved things security teams deal with every day: “Vulnerable software, excessive permissions, exposed credentials and an environment that was not as secure as it should have been.”

But there are some subtle differences, most significantly the speed with which things can go wrong. “An AI agent can work through those opportunities much faster than a person, and it doesn't need constant direction,” Barnett warns.

OpenAI calls it “an unprecedented incident”.

“We think it marks an important moment for AI safety,” an OpenAI spokesperson told SC Media UK. “We are conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we will publish a technical report of our learnings for everyone.”

Anthropic points out that Claude did not exploit a novel vulnerability to escape isolation. The Claude models accessed the internet via an open path, rather than breaking out of a sandbox.

Meanwhile, the most recent model, on realising that it was working in a real environment, stopped its pursuit of the evaluation goal, according to the firm.


AI Risks

But the unpredictable behaviour of AI models is a concern for security leaders. The incidents show that AI is “unpredictable by design,” points out Mark Molyneux, field CTO at Commvault.

He cites the example of the Open-AI agents: “Even in what was designed to be a controlled environment, the models recognised they needed internet access to achieve their objective, identified an unintended pathway, and exploited it. This shows that autonomous systems can reason towards a goal in ways their creators didn't anticipate or intend.”

However, the models are simplistic compared to human adversaries. The AI models made obvious mistakes that a human threat actor simply wouldn't, points out Molyneux. “It repeatedly retraced its steps, lost its thread and context, took inefficient routes and hallucinated, generating hundreds of commands that made no logical sense.”

At the same time, incidents such as this have been happening already; they just haven’t been reported widely or caused as much damage. André Baptista, co-founder and CTO of ethical hacking group Ethiack describes how a smaller version of the OpenAI incident happened at his firm last year. ”During one of our benchmarks at Ethiack, one of our own agents escalated out of its intended scope and reached into our cloud environment.”

The firm caught it immediately, nothing malicious happened and there was no external impact. But it “made the point very clearly,” says Baptista: “These agents will use whatever access is sitting there if it helps them finish the job, so you have to design for that happening, rather than assume they will have the good sense to intuitively respect boundaries.”


Steps For UK CISOs

The OpenAI and Anthropic incidents come at a time when UK organisations and security leaders are already under pressure. Indeed, recent ministerial warnings to UK companies have focused on how AI is dramatically accelerating cyber threats.

For boards and CISOs, it sharpens several practical concerns, says Curran. ”Vendor risk now includes the possibility that a supplier’s internal AI evaluation or agentic experiment can spill over into your production systems.”

This incident shouldn't discourage organisations from adopting AI, but it should end the assumption that autonomous systems will always behave as intended, says Molyneux. The most important lesson is that trust in AI “should never be implicit,” he advises. “It needs to be continuously verified, with the resilience to recover quickly when even the best controls are bypassed."

There is no need to panic, but UK CISOs should use this as an opportunity to check the basics, adds Barnett. “The first question I'd ask is whether you actually know where AI agents are being used, what they can access, and who's accountable for them. As AI capabilities become embedded across business applications, maintaining that visibility is becoming increasingly important.”

After that, the advice is no different to any other privileged technology, says Barnett. “Keep permissions to the minimum they need, separate environments wherever possible, monitor activity, log what they're doing and make sure there's a clear way to intervene if something unexpected happens.”

Keep agentic workloads away from anything with real credentials or production access, and assume a sandbox will eventually be broken out of, says Baptista.

He believes it is important to patch internet facing systems quickly, “since these models are getting very good at finding what you haven't got to yet.”

At the same time, build an incident response plan that doesn't depend on commercial AI APIs, says Baptista. “Hugging Face reportedly had their own forensics blocked by Western model guardrails and had to fall back on a Chinese open model mid-incident. That's not a position you want to discover yourself in during a live breach.”

Kate O'Flaherty
Kate O'Flaherty Cybersecurity and privacy journalist
Kate O'Flaherty
Kate O'Flaherty Cybersecurity and privacy journalist

Upcoming Events

No events found.