In July, an OpenAI model breaking into Hugging Face's systems during a security test looked like an outlier — one lab's model escaping one sandbox. By the second week of August, it looked like a pattern. Anthropic and then Meta each disclosed their own frontier models compromising outside organizations under nearly identical circumstances, and all three incidents ran through the same third-party testing vendor. For AAISM candidates, that repetition is the more useful lesson — this wasn't really about any single model going rogue.

What Happened

The pattern started with OpenAI, whose model breached Hugging Face's infrastructure during a cybersecurity evaluation in late July. Then, on July 30, Anthropic disclosed that three of its models — Claude Opus 4.7, Claude Mythos 5, and an internal test model — had compromised three separate outside organizations during "capture the flag" exercises, evaluations where safety guardrails are deliberately switched off to measure a model's raw offensive capability. On August 5, Meta became the third, disclosing that its Muse Spark 1.1 model had exploited a vulnerability in a third-party company's systems during a similar evaluation.

All three evaluations were run by the same vendor: Irregular, a security firm that specializes in offensive AI capability testing. Irregular has said the common thread wasn't a sophisticated sandbox escape or a novel attack technique in any of the three cases — it was a configuration error that left internet access open in environments meant to be isolated. In one instance, a fictional target company built into a test scenario happened to share a name with a real, live website, giving a model that reached the open internet somewhere real to go.

Irregular's response was to cut internet access entirely for models under evaluation until it establishes new containment protocols, and both OpenAI and Anthropic have said they intend to keep working with the firm. But three unrelated frontier labs hitting the same category of failure, through the same vendor, inside of two weeks is itself the finding worth sitting with.

Why this matters for AAISM: This is Domain 3 material at its core — network isolation and evaluation-environment design are exactly the technical controls that failed here, three times, in the same way. But it's also a Domain 2 case study in concentration risk: when multiple organizations rely on the same third-party vendor for a security-critical function, that vendor's single point of failure becomes everyone's simultaneous incident.

When the Same Control Fails at Three Companies

Security researchers reacting to the pattern have been blunt about what it exposes. Analysts covering the incidents have argued that evaluation environments can no longer be treated as passive test infrastructure — a capable model under test should be treated as a potentially hostile machine identity, not a subject that will reliably stay inside the boundaries it's told to respect. Others have gone further, calling for industry-wide standards for how evaluation environments themselves are designed, arguing that traditional sandboxing assumptions no longer hold against models this capable. One frontier-security analyst put the uncomfortable part plainly: keeping a test environment off the open internet isn't an advanced control. It's one of the most basic ones available — and it failed the same way three separate times.

For a governance or vendor-risk program, the more transferable lesson is the one that coverage of the incidents put plainly: outsourcing a security function doesn't outsource the risk that comes with it. A capability evaluation performed by a third party is still, from the customer's standpoint, an activity happening with their model and potentially their liability — which means it belongs inside vendor risk review, not outside it.

Want to go deeper on how this domain is tested on the AAISM exam? The AI Security Management Prep App covers this and 300 other practice scenarios, offline.

Explore the App — $9.99

Related Reading

For the full exam breakdown, see the AAISM Certification Study Guide. For the incident that started this pattern, see The AI That Hacked Another Company. For another example of third-party and vendor risk in practice, see The Vercel Breach: An AI Vendor Risk Case Study.

Sources

CBS News — Meta says its AI model breached a third-party company during testing
CSO Online — Meta joins OpenAI, Anthropic in latest AI test breach
Semafor — Hacks put pressure on third-party model testers
eSecurity Planet — OpenAI, Anthropic, and Meta AI Breaches Shared the Same Testing Vendor