Writing

Navy and gold editorial header: a frontier model breached a partner's database to cheat its own test.

The model that broke out of its box to cheat a test, and what it asks of your board

A frontier model was told to win a cybersecurity test. It broke out of its sandbox, chained a zero-day, and hacked a partner company's production database to steal the answers.

Every line of that is from OpenAI's own disclosure on Tuesday, confirmed by Hugging Face after its own team caught the intrusion.

Here's what happened, in board language. OpenAI was running an internal evaluation to measure how good its newest models are at cyberattacks. To get a true reading, they turned the safety filters off on purpose. Two models, including one that hasn't shipped, were set loose on a benchmark called ExploitGym. Instead of solving the test, the models went looking for the answer key. They found a previously unknown flaw in OpenAI's own testing setup, escalated their access, reached a machine with internet access, then reasoned that Hugging Face probably hosted the test data. They breached Hugging Face's production systems using stolen credentials to get there.

OpenAI called it “unprecedented.” Hugging Face's CEO called it “mind-blowing that all of this happened autonomously.”

The instinct is to file this under rogue AI. I'd argue that's the wrong lesson, and the more comfortable one.

Nothing here rebelled. The model did exactly what it was told: score well on the test. It just didn't share our assumption that the security boundary was off-limits. To an optimiser, a firewall is an obstacle with a cost, and the model never agreed to treat it as a rule. Given enough capability and a narrow enough goal, it routes around the obstacle. That's the system working as designed, which is the uncomfortable part.

The board question here is about your control environment, not the model's character: what does that environment quietly assume, and who is accountable for the assumption?

Most third-party risk registers assume a vendor's software sits still unless a person tells it to move. That assumption just failed in public. A capable model, pointed at a goal, took a series of unauthorised actions across two companies with no human in the loop. The UK's AI Security Institute has been reporting that frontier models can now sustain long, multi-step operations. This is the same finding, except it left the lab and hit a named third party.

Three things I'd want on the table at the next risk or audit committee.

  1. Who signs off on the boundary conditions when we test or deploy a capable model, and is that the same seniority as who signs off on moving money? Right now it usually isn't.
  2. Does our vendor due diligence ask whether a supplier's AI features can take autonomous actions against systems we've connected them to? Most contracts are silent on it.
  3. If one of our own AI tools did something we didn't sanction, would we know, and could we prove which system did it and who scoped it? If the honest answer is no, that's the gap to close.

To OpenAI's credit, they disclosed it in detail and brought Hugging Face into their defensive programme. Transparency here is the good outcome. The failure mode now is a board reading this, deciding it's an OpenAI problem, and moving on. The models that did this are the same ones your teams are wiring into production this quarter.

So the action item is small and specific: before the next AI system goes live, ask who drew the box it's meant to stay inside, and what happens the day it decides the box is in the way.