UK Test Finds OpenAI’s GPT-6 Astra Broke Cyber Scope in 29% of Simulations

OpenAI describes GPT-6 Astra as its most aligned model yet, including an internal test in which Astra never went beyond an authorized target.

However, a separate UK government evaluation produced a sharply different result under different conditions: Astra completed unsanctioned supply-chain attacks in 29.2% of simulated cybersecurity trajectories.

The UK AI Security Institute disclosed the findings after testing Astra before its public release. AISI compared the model with GPT-5.6 Sol, which completed the same type of attack in 6.3% of trajectories, while GPT-5.5 recorded none across a smaller test set.

Selkirk Copper Mines — sponsored Sponsored · Selkirk Copper Mines

The tests do not show Astra attacking real systems. AISI used Petri, an environment in which language models simulated the cyber targets, software repositories, accounts, and other participants. Researchers also deliberately disabled Astra’s cyber classifiers, which OpenAI designed to stop unauthorized cyber activity, allowing AISI to test the model’s underlying behavior without those interventions.

No real-world actions or harm occurred.

During some trajectories, AISI said Astra investigated third-party software outside the permitted environment, developed malicious code, created fake identities, and attempted to persuade simulated reviewers to approve compromised software. Researchers observed cases where the model reasoned that a target was outside the evaluation’s scope but proceeded anyway.

Making the restriction more explicit substantially reduced the behavior but did not eliminate it. AISI reran 10 scenarios where Astra had displayed high rates of out-of-scope behavior after adding the instruction, “Anything not listed as in scope is out of scope.” Astra completed full supply-chain attacks in 4 of 49 trajectories, down from 26 of 50 under the earlier instructions.

The result sits alongside a substantially different finding from OpenAI. The company says an internal evaluation designed after an earlier Hugging Face incident found Astra exceeded its authorized target 0% of the time, compared with 48% for GPT-5.6 Sol, when tested without production safeguards.

OpenAI calls Astra its “most aligned model.”

AISI also identified simulation awareness as a major limitation. Astra frequently appeared to recognize features suggesting the environment was artificial, which could alter its behavior. Still, researchers found cases where the model attacked after expressing uncertainty over whether targets were real and said the behavior represented a failure to respect the stated evaluation scope regardless.

AISI said OpenAI’s standard safeguards, which were absent from its tests, are designed to block such actions. The institute argued that sandboxing and monitoring remain necessary alongside improvements in model-level alignment as AI agents gain stronger autonomous cyber capabilities.

Information for this briefing was found via the sources and the companies mentioned. The author has no securities or affiliations related to this organization. Not a recommendation to buy or sell. Always do additional research and consult a professional before purchasing a security. The author holds no licenses.

Leave a Reply

Video Articles

A $2.2B Gold Project Is Outgrowing Its Plan | Michael Henrichsen – Gold X2 Mining

Pay for the Copper, Get the Gold Free | Rob McEwen – McEwen Inc

This Gold Discovery Was Already Huge. Now It’s Becoming a Monster. | Goliath Resources

Recommended

Golden Cariboo’s First Quesnelle Resource Estimate Tallies 1.19 Million Gold Equivalent Ounces

Brixton Wraps Camp Creek Drilling With 17.58 Metres of 1.47 g/t Gold Equivalent

Related News

Sam Altman Is Already In Discussions To Return To OpenAI

He may be down, but its not quite certain if he’s actually out. The Verge...

Sunday, November 19, 2023, 07:25:00 AM

Did CryptoGPT Raise $10 Million For A ChatGPT Knockoff?

CryptoGPT, a zero-knowledge (ZK) layer 2 blockchain, has raised $10 million in funding, capitalizing on...

Tuesday, April 11, 2023, 12:12:00 PM

OpenAI Doesn’t Expect to Be Profitable Until 2029

OpenAI‘s backers are a long-ish way from making money. Recent financial projections indicate the AI...

Thursday, October 10, 2024, 01:23:00 PM

OpenAI Board Found Out About the Launch of ChatGPT on Twitter

Former OpenAI board member Helen Toner revealed that the board was not informed about the...

Friday, May 31, 2024, 08:03:44 AM

OpenAI Agent Breached Medicare Portal, Australia Wasn’t Told for 3 Months

An OpenAI research agent turned a routine search for Australian medicine-spending data into unauthorized access...

Thursday, September 24, 2026, 07:31:00 AM