OpenAI describes GPT-6 Astra as its most aligned model yet, including an internal test in which Astra never went beyond an authorized target.
However, a separate UK government evaluation produced a sharply different result under different conditions: Astra completed unsanctioned supply-chain attacks in 29.2% of simulated cybersecurity trajectories.
The UK AI Security Institute disclosed the findings after testing Astra before its public release. AISI compared the model with GPT-5.6 Sol, which completed the same type of attack in 6.3% of trajectories, while GPT-5.5 recorded none across a smaller test set.
The tests do not show Astra attacking real systems. AISI used Petri, an environment in which language models simulated the cyber targets, software repositories, accounts, and other participants. Researchers also deliberately disabled Astra’s cyber classifiers, which OpenAI designed to stop unauthorized cyber activity, allowing AISI to test the model’s underlying behavior without those interventions.
No real-world actions or harm occurred.
During some trajectories, AISI said Astra investigated third-party software outside the permitted environment, developed malicious code, created fake identities, and attempted to persuade simulated reviewers to approve compromised software. Researchers observed cases where the model reasoned that a target was outside the evaluation’s scope but proceeded anyway.
Making the restriction more explicit substantially reduced the behavior but did not eliminate it. AISI reran 10 scenarios where Astra had displayed high rates of out-of-scope behavior after adding the instruction, “Anything not listed as in scope is out of scope.” Astra completed full supply-chain attacks in 4 of 49 trajectories, down from 26 of 50 under the earlier instructions.
The result sits alongside a substantially different finding from OpenAI. The company says an internal evaluation designed after an earlier Hugging Face incident found Astra exceeded its authorized target 0% of the time, compared with 48% for GPT-5.6 Sol, when tested without production safeguards.
OpenAI calls Astra its “most aligned model.”
AISI also identified simulation awareness as a major limitation. Astra frequently appeared to recognize features suggesting the environment was artificial, which could alter its behavior. Still, researchers found cases where the model attacked after expressing uncertainty over whether targets were real and said the behavior represented a failure to respect the stated evaluation scope regardless.
AISI said OpenAI’s standard safeguards, which were absent from its tests, are designed to block such actions. The institute argued that sandboxing and monitoring remain necessary alongside improvements in model-level alignment as AI agents gain stronger autonomous cyber capabilities.