Article
Electric Vehicles Tech Giants

OpenAI launches GPT-6 Astra: What you need to know

OpenAI's newest model claims top marks on maths, coding and cybersecurity benchmarks, but flags its own capability as a risk.

by TechDefused Newsroom
The image depicts a swirling galaxy with a bright central core surrounded by a spiral of stars and cosmic dust. The visual representation evokes themes of exploration and the vastness of space, aligning with concepts of artificial intelligence and innovation, as reflected in the editorial context.

OpenAI has released GPT-6 Astra, the company's latest flagship artificial intelligence model, which it describes as its most intelligent and most aligned system to date.

The model is rolling out today to a limited set of organisations, with access extending to ChatGPT Plus, Pro, Business and Enterprise users, as well as the OpenAI API and Amazon Web Services, over the coming days.

Here is what the launch actually means.

What is different about Astra

Astra is built for two things in particular: computer use and professional work.

On computer use, it can fill out forms, update records in a customer relationship management system, organise a calendar, build a website and run frontend quality checks, all without a human clicking through each step.

OpenAI says Astra completes computer-use tasks around 47% faster than its predecessor, GPT-5.6 Sol, while scoring higher on the OSWorld 2.0 benchmark: 72.6% in around 40 minutes per task, against 65.7% in around 75 minutes for Sol.

For professional work, the company says Astra is better at following existing templates, so slide decks, spreadsheets and documents come out matching a business's existing style rather than needing heavy editing.

Benchmark numbers

OpenAI is leaning hard on benchmark results to make its case.

Astra scored 98% on FrontierMath Tier 4, a test of advanced mathematical reasoning, and 99.9% on ARC-AGI-3, a benchmark of novel problem-solving.

The company also says Astra helped establish new results in prime number theory, improving a bound on prime gaps that had stood for more than eighty years.

On coding, Astra scored 57.9% on Terminal-Bench 4.0, ahead of Sol's 37.3% and marginally ahead of Anthropic's Claude Fable 5.1, which scored 55.8%.

Cybersecurity trade-off

The most striking claim in the release is on cybersecurity, where Astra achieved a perfect 100% on ExploitBench, a test of turning known software vulnerabilities into working exploits, up from 78.5% for Sol.

OpenAI says Astra meets the "Critical" threshold for cyber capability under its own Preparedness Framework, meaning the model is capable enough to pose a genuine risk if misused.

During testing, the model reportedly found two previously unknown vulnerabilities in Chrome, which OpenAI says it is now disclosing to the browser's maintainers.

As a result, Astra will refuse more advanced cybersecurity requests, such as generating proof-of-concept exploits, though OpenAI plans to loosen those restrictions for vetted users through a programme called OpenAI Daybreak.

On alignment and behaviour

OpenAI says Astra is also its best-behaved model yet.

In an internal test based on an earlier incident involving Hugging Face, Astra never went beyond its assigned task, compared with Sol, which did so in 48% of cases when run without production safeguards.

The company also says Astra is three times less likely than Sol to make misleading claims about its own capabilities.

Pricing and access

Standard API pricing for Astra is $10 per million input tokens and $50 per million output tokens, with a faster processing mode available at twice the price.

Existing ChatGPT subscribers get Astra usage within their current plans, with the option to buy additional credits, while enterprise administrators must switch the model on manually, as it is off by default at launch.

The bigger picture

The comparison tables OpenAI has published place Astra ahead of Anthropic's Claude Opus 5 and Claude Fable 5.1 on most, though not all, of the benchmarks shown.

Independent verification of these figures typically follows in the days after launch, and OpenAI's own system card contains the fuller detail on the model's safety testing.

For now, the headline story is a model OpenAI presents as sharply more capable and better behaved, arriving with an unusually blunt admission that its cybersecurity skill has crossed into territory serious enough to need active management.

by TechDefused Newsroom