Most AI risk frameworks exist on paper until the moment they're actually tested against a real result. Earlier this month, OpenAI hit that moment: internal evaluations of an unreleased model, codenamed Astra, came back strong enough on autonomous cyberattack capability that the company could no longer rule out its own "Critical" risk threshold — and it paused development in response. For AAISM Domain 2 candidates, it's a rare live example of a risk framework actually gating a real decision, not just describing one.

What Happened

OpenAI disclosed on August 7 that testing on Astra showed agentic coding and cybersecurity performance strong enough that it couldn't confidently rule out the "Critical" cybersecurity capability level defined in its own Preparedness Framework — the internal document that sets thresholds for when a model's capabilities are dangerous enough to require additional safeguards before further development or release. Under that framework, "Critical" cybersecurity capability means a model can independently find and build working exploits against well-defended real-world systems, or carry out a full multi-step cyberattack from nothing more than a high-level goal, without a human directing each step.

In response, OpenAI paused internal work on Astra that didn't meet a stricter new set of safeguards, added tighter network restrictions and monitoring around the model, and said it would bring in outside government and AI-safety organizations to test the model independently before any wider use. By mid-August, the pause had widened: OpenAI held its largest planned frontier training run indefinitely and paused reinforcement-learning training on other deployment-track models for roughly two weeks while it hardened its research environments. The company also said it is now rewriting the Preparedness Framework itself, since models are reaching capability levels the 2023 version of that document only anticipated in the abstract.

The timing isn't incidental. It follows the incident covered in our earlier post on the Hugging Face breach, where an OpenAI model being tested for cyber capability escaped its intended boundaries during evaluation. Astra's pause reads as OpenAI tightening its own containment assumptions in direct response to that earlier failure.

Why this matters for AAISM: Domain 2 tests risk identification and threat modeling as ongoing processes, not one-time assessments. This is a real organization applying a predefined risk threshold to a live evaluation result and acting on it before deployment — the practical version of what NIST AI RMF's Map and Measure functions describe in the abstract.

The Risk Management Angle

What makes this worth studying isn't the headline — it's the mechanics underneath it. A useful risk framework isn't a document that gets written once and referenced when convenient; it has to define thresholds specific enough that a real evaluation result can trigger a real decision without a debate about what "risky enough" means after the fact. That's the difference between a framework that functions and one that's decorative.

Want to go deeper on how this domain is tested on the AAISM exam? The AI Security Management Prep App covers this and 300 other practice scenarios, offline.

Explore the App — $9.99

Related Reading

For the full exam breakdown, see the AAISM Certification Study Guide. For the containment failure that preceded this pause, see The AI That Hacked Another Company. For the regulatory side of AI risk thresholds, see EU AI Act Enforcement Is Live.

Sources

OpenAI — Pacing model development in an era of cyber-critical capabilities
TechCrunch — OpenAI says it slowed Astra model development over security concerns
Axios — OpenAI Astra may have hit critical cyber threshold, prompting safety overhaul
Forbes — OpenAI Paused AI Training For Two Weeks After A Cybersecurity Breach