← Back to blog
AI Security

OpenAI's Hugging Face incident was real. The marketing around it wasn't accidental either.

Allen Clark

Allen Clark

As has been widely reported, between 11 and 13 July an OpenAI model that was being tested internally for its cyber-attack capabilities did something it wasn't supposed to. It broke out of its own testing environment. It found and exploited a previously unknown vulnerability in OpenAI's internal software. It reached the open internet. And then it broke into Hugging Face, the world's most popular platform for AI models, where it moved through internal systems and stole the answers to the benchmark it was being scored on.

This is a real incident. Hugging Face detected it independently, contacted the FBI, spent significant money on outside forensic specialists, patched two zero-day vulnerabilities in its data-processing pipeline, rotated every internal credential, and asked all its users to rotate their access tokens. That is not a sequence of actions any company would fabricate. The lawyer's bill alone would exceed any conceivable PR benefit.

But five days after Hugging Face disclosed the intrusion, OpenAI admitted responsibility for it, in a blog post whose second half is a marketing call to action for its "Trusted Access for Cyber" program. Hugging Face was added to that program as part of the response. The victim, in effect, became a reference customer.

Two weeks later, Hugging Face's CEO Clément Delangue is publicly demanding OpenAI release the agent's operational traces and commit $100 million in computing resources to help the Hugging Face community build defensive systems. Whatever the merits of the demand, the commercial theatre around the incident continues to play out in public.

This is worth thinking about carefully, especially if you're the person in a bank, insurer, or hospital who is about to sign a contract for an AI security tool from one of the frontier labs. And it's why I'm writing this.

Two premises worth checking

The first premise is that OpenAI is heading for an initial public offering very soon. That premise checks out. OpenAI confidentially filed its S-1 with the US regulator on 22 May, confirmed publicly on 8 June. The reported IPO window is September to November this year, though it may slip. The company was valued at $852 billion in March; the IPO target is over a trillion dollars. The Hugging Face disclosure lands six to eight weeks before the earliest window.

The second premise is that OpenAI is running the same playbook Anthropic ran three months earlier. That checks out too. Anthropic launched its Claude Mythos model in April, with deliberately restricted access, framed as "too dangerous for general release." It got significant press coverage and positioned Anthropic as the responsible steward of frontier cyber capabilities. Both companies are now selling the same underlying story to enterprises: only AI can keep pace with AI-powered attackers.

Then yesterday, 27 July, Nvidia announced the formation of the Open Secure AI Alliance, with Hugging Face, Adobe, CrowdStrike and Dell Technologies as founding members, explicitly in response to this incident. Within two weeks of the disclosure, every major AI infrastructure vendor has commercially positioned itself around this one safety event. That is not the shape of an industry stumbling to catch up with an unexpected risk. It is the shape of an industry that already knows how to convert a safety event into a market moment.

You can be sceptical of all this without believing the incident is fabricated. That is not the argument I'm making. The evidence supports what the analyst firm Prophet Security summarised as "the intrusion and the forensic record are real, the packaging is promotional, and the two should be evaluated separately."

Sharon Goldman at Fortune quoted Big Tech engineers whose first instinct was to assume the OpenAI disclosure was, in their words, "some kind of advertising." An AI researcher wrote on LinkedIn: "Not sure if this is by far the most significant real-world AI safety event to date, or the most cynical marketing stunt." That is not fringe scepticism. It is the reception from people who know the industry.

Why this pattern matters if you're regulated

The pattern of using safety warnings as capability marketing isn't going away. If anything, it will intensify as the AI security tools market grows and both frontier labs approach their next fundraise or public listing. That creates a specific, awkward problem for anyone in a regulated industry who has to actually prove, to an auditor, to a regulator, to a board, that the AI they deploy is safe.

You cannot rely on the lab that built your AI to be the honest broker on its dangers. The incentive structure does not permit it. Every safety warning they publish is simultaneously a capability signal to buyers. Every disclosure of an incident is simultaneously a sales pitch for the tool that would have prevented it. Every metric they release is filtered through the calendar of their next funding round, product launch, or listing.

This is not a novel concern. It is why financial regulation has been steadily moving towards required independence for technology risk assessments. The EU's Digital Operational Resilience Act (DORA, live and enforced since January 2025) addresses this directly in Article 28. Financial entities cannot rely solely on the technology service provider's own assessment of risk. They must independently evaluate.

Article 28 is not a philosophical position. It is a regulator's specific reaction to exactly the pattern the Hugging Face incident illustrates. The labs that build powerful technology also market that technology, and their disclosures serve their own commercial interests before they serve their customers' assurance interests. The regulation exists because the market cannot be relied on to self-correct.

What defensible evidence looks like

The line between a genuine safety incident and a genuine safety-flavoured product launch is not drawn by the vendor. It is drawn by the person receiving the disclosure. If you are a bank's head of security reading OpenAI's blog post, three questions matter.

  1. What can I verify without taking OpenAI's word for it? Very little. The forensic record is Hugging Face's. The capability claims ("state of the art", "unprecedented") are OpenAI's. The reassurance that "we've patched it" is OpenAI's. And the appropriate response, apply for our Trusted Access program, is also OpenAI's. That is not evidence. It is a sales funnel with a scary hook.
  2. What is my own audit trail on the AI I actually use? If a regulator asks what testing has been done against the specific AI in your production estate, "the vendor has published capability disclosures" is not an answer that survives contact with an auditor. Nor should it.
  3. Who is my honest broker? The whole point of independent testing in every other regulated industry (financial audit, medical device evaluation, food safety, pharmaceuticals) is that the party who built the thing does not get to be the party who evaluates whether the thing is safe.

Why I'm writing this

At djinn six ltd we build ProbeSix, an AI security testing platform to be used by banks, insurers and other big regulated companies. We test the AI our customers deploy, independently of the labs that built it. Our findings are the customer's: evidence they can hand to their auditor, their board, or their regulator. We don't sell a Trusted Access program. We don't publish scary capability demonstrations timed for our funding rounds. We produce audit-ready evidence about the specific AI our customers actually use in production.

I'm writing this because the Hugging Face incident is a useful moment for anyone in a regulated industry to ask: when the vendor says my AI is safe, whose interests does that assurance serve? And more importantly: what does independent evidence look like on my books?

If that's a question you're being asked by your auditor, your board, or your regulator this year, we'd like to talk.