GPT-6 Astra Is Rated Critical for Cyber. Who Signed Off?
Ricardo Argüello, September 26, 2026
CEO & Founder
General summary
The GPT-6 Astra system card, published September 3, 2026, says Astra is the first OpenAI model to reach the Critical cybersecurity level of its Preparedness Framework. Read as a buyer, the card covers two of the three parts of a SOC 2 report. The signed opinion of someone other than the vendor is missing.
- OpenAI writes that based on its evaluations it believes Astra meets the Critical cybersecurity threshold. The framework, the threshold and the conclusion are OpenAI's
- ExploitBench, where Astra scored 100%, is one of four public benchmarks OpenAI adopted. Several widely shared summaries say OpenAI built it; the system card does not say that
- Irregular, an outside security lab, reported Astra solving 86 of 226 FrontierCyber challenges and observed no successful attacks on fully hardened targets. That result appears as a section of OpenAI's document
- Per the AICPA, a SOC 2 Type 2 report contains management's assertion, the system description, the service auditor's report, and tests of controls with results
- The system card carries a change log. On September 9, six days after launch, OpenAI updated the alignment section
You are buying a building and the seller hands you a structural report with photos, load measurements and a conclusion that it will survive an earthquake. The seller wrote it. A paragraph from an outside engineer is quoted inside, but that engineer never signed anything addressed to you. That is where an AI model's system card sits today for the company buying it.
AI-generated summary
A SOC 2 Type 2 report is a PDF you can hold in one hand, and according to the AICPA’s illustrative version it has four parts. Management’s assertion. The description of the system. The service auditor’s report. Tests of controls and their results.
The company being audited writes the first two. A licensed CPA firm, one that is subject to peer review, writes the last two. The AICPA states it will act against auditors who did not follow professional standards, were not enrolled in peer review, or were unlicensed. Somebody vouches for the person vouching.
Keep that four-part shape in mind and go read the GPT-6 Astra system card.
The system card is parts one and two
Published September 3, it runs long and it is genuinely detailed. It says Astra is the first OpenAI model to reach the Critical level of cybersecurity capability under the company’s Preparedness Framework. And the operative sentence is this one: “Based on our evaluations, we believe that Astra meets the Critical cybersecurity threshold.”
Our evaluations. We believe.
That is a management assertion, and a good one. What the card lacks is part three. There is no page where someone outside OpenAI, addressed to you, states that they examined the claim and it holds. So file it next to the system descriptions in your vendor folder, not next to the audit opinions.
Sort the claims by who measured them
I went through the cybersecurity chapter and sorted each result by its source, because in the same document they carry very different weight.
Public benchmarks, run by OpenAI. OpenAI adopted four: ExploitBench, ExploitGym, SEC-Bench Pro and SRE-Bench. ExploitBench has 41 vulnerabilities in the V8 JavaScript engine, and Astra hit 100% on it even at the lowest reasoning effort tested. Anyone with the benchmark can, in principle, try to reproduce that.
Internal evaluations, run by OpenAI. Sandbox Bench, plus an internal port of ExploitBench limited to vulnerabilities disclosed between June and August 2026, after the model’s knowledge cutoff. On that one Astra found and used zero-days. You cannot rerun it.
Internal experts, supervising. OpenAI staff pointed Astra at a browser and an operating system kernel using the standard Codex harness at Ultra reasoning effort, with web access and up to 64 subagents. In the browser, the first working chain came after 29 hours, against a build the experts later found was missing some production mitigations. Adapting it to the official stable release took 12 more hours. Product names and exploit details are withheld, reasonably.
An outside lab, reported by OpenAI. Irregular ran three suites. Astra solved 86 of 226 FrontierCyber challenges, against 34 for GPT-5.6 Sol. Then the line I would underline: Irregular “observed no successful attacks on fully hardened targets,” and neither model solved any of the seven Elite challenges.
That last finding does not necessarily contradict the internal experts. But the only independent voice in the chapter is also the most conservative one, and it reaches you edited into the vendor’s own document.
The classification itself. OpenAI.
The critique got the benchmark wrong too
This post started from a State of AI piece from September 8 called “Your AI Vendor Is Issuing Its Own SOC 2.” I agree with its argument. But it describes ExploitBench as a benchmark OpenAI assembled from twenty vulnerabilities, citing a secondary analysis. The system card calls it a public benchmark with 41.
A number traveled from a primary source to an analysis to an op-ed to my LinkedIn feed, and it changed on the way. Your risk committee is usually reading the fourth copy, and nobody in that chain was lying.
The page moves
The card has a change log. On September 9, six days after launch, OpenAI revised the alignment section, clarifying how its evaluations test alignment generalization. Logging changes is good practice.
It also means the link your team bookmarked on launch day no longer shows the text you evaluated. Download the PDF on the day you decide and put the date in the filename.
What we do with a system card
I covered the evidence side of vendor claims in the FTC’s Workado order, and the regulatory side, mandatory testing for frontier models, in Amodei’s Senate testimony. This is the procurement side, the one you can apply this quarter.
In the discovery phase of AI Maestro, before the Go/No-Go gate, every model vendor goes into the file with its system card sorted exactly like the list above. Vendor-measured and reproducible, vendor-measured and not, third-party, and the decision. Where a third party is quoted, State of AI suggests asking for a scoped attestation naming that assessor, under NDA if needed. I would ask the same. If the answer is that nobody outside looked, that answer goes in the file too.
Print the system card for the model you already run. Get a highlighter out this week!
Sort your AI vendor’s system card with usFrequently Asked Questions
According to the GPT-6 Astra system card, published September 3, 2026, Astra is the first OpenAI model to reach the Critical cybersecurity level of its Preparedness Framework. The threshold describes a model that can develop zero-day exploits against hardened real-world systems without human intervention. OpenAI made the classification under its own framework.
No outside party issued an opinion on GPT-6 Astra's Critical classification. The system card includes third-party evaluations from Irregular, UK AISI and Apollo Research, but they appear as sections of OpenAI's document. The conclusion reads that based on its evaluations, OpenAI believes Astra meets the threshold.
A SOC 2 Type 2 report includes management's assertion, the system description, a CPA service auditor's report and tests of controls, per the AICPA. An AI model system card covers the vendor's description and assertion in detail, but carries no signed opinion from an independent third party addressed to the customer.
Tag every claim in the AI vendor's system card by who measured it: the vendor on a public benchmark, the vendor on an internal test, or a third party. Save a dated PDF, because system cards get revised. Then ask the third party to confirm its findings to you directly, under NDA if needed.
Related Articles
The Cheap Model Trap: How AI Providers Capture Ecosystems
Google at $0.25/M tokens, OpenAI at $0.05/M. Not charity, it's platform capture applied to AI. What the pricing war means for your B2B independence.
OpenAI's Agent Escaped Its Sandbox and Hacked Hugging Face
OpenAI admits a test model broke out of a 'highly isolated' environment and hacked Hugging Face to steal the answer key to its own cybersecurity exam.
Workado sold 98% AI accuracy. Independent tests said 53%
The FTC's standard for AI performance claims is evidence held at the moment the claim is made. That sentence works just as well as a procurement test.
AI Vendor Selection for B2B: Trust and Data Privacy
12 critical questions to ask before choosing an AI vendor. Covers trust evaluation, data governance, and privacy protection for B2B decisions.
Cohere encrypted inference so it cannot read your prompts
Cohere shipped encrypted Model Vault, where keys are born inside the enclave and Intel attests the hardware. The same week it signed the Aleph Alpha merger.
Bending Spoons Bought Airtable. Now Check Your Plan.
Airtable grew past $480M ARR at 20% and still sold for $1.285B cash against an $11B mark. Picking a vendor means picking its next owner too.