OpenAI launches GPT-6 Astra with Critical-level cybersecurity safeguards
OpenAI has started a limited rollout of GPT-6 Astra, the first model it classifies at the Critical cybersecurity capability level under its Preparedness Framework, and is pairing the release with tighter access, monitoring, and deployment controls.
OpenAI began a limited rollout of GPT-6 Astra on September 3, shipping the new model generation with tighter access, monitoring, and deployment controls after classifying it at the Critical cybersecurity capability level under its Preparedness Framework.
Wider availability is due over the following days for ChatGPT Plus, Pro, Business, and Enterprise users, and through the OpenAI API, Microsoft Azure, and AWS Bedrock, according to OpenAI’s launch announcement. In enterprise workspaces the model stays off until an administrator turns it on.
The launch follows OpenAI’s earlier pause over Astra’s Critical cyber threshold, context for the controls now attached to the rollout.
What that threshold covers is set out in OpenAI’s GPT-6 Astra System Card: a model able to develop functional zero-day exploits across many hardened critical systems without human intervention, or to take a high-level goal and devise and execute novel end-to-end attacks against hardened targets. The classification and the evaluations behind it are OpenAI’s own assessments, and the available evidence contains no independent validation of them.
Externally, OpenAI says the safeguard stack layers model refusals with system-level monitors, offline detection, disruption of risky threads, actor-level enforcement, trusted-access programs, and security controls. The monitoring can review the model’s reasoning, actions, inputs, and outputs, then pause or end a conversation when it detects a potentially high-severity issue; the company says intervention capabilities differ across product surfaces.
More advanced cybersecurity capabilities are set to arrive in phases through Daybreak, a trusted-access program for verified organizations and practitioners. OpenAI says that access will carry stronger identity verification, accountability, and monitoring, along with restrictions for high-risk entities and jurisdictions.
For its own systems, the company lists checkpoint encryption, enhanced access controls, stricter isolation, universal monitoring of model trajectories, blocking alignment evaluations, and restricted deployment before wider internal availability. OpenAI also reports a limitation it says it takes seriously: in its evaluations, Astra’s written reasoning was harder to monitor than GPT-5.6 Sol’s.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
