OpenAI Says GPT-6 Astra Opens the AGI Era. The Benchmarks Tell a More Nuanced Story

“Welcome to the AGI era.” That is how OpenAI president Greg Brockman greeted the release of GPT-6 Astra, the model the company has spent months teasing as its biggest launch of the year.

It is a striking claim, and Brockman doubled down when asked directly whether Astra qualifies as artificial general intelligence. “For me personally, I do think we’re there.” Sam Altman’s framing was slightly more measured, calling it a generational leap rather than a finish line.

So has the industry’s most contested label finally been earned? The early numbers make the case for OpenAI, and against it, in equal measure.

The scores that made people sit up

OpenAI describes Astra as the most intelligent and aligned model in the world, with new benchmark highs across science, mathematics, computer use, coding and cybersecurity.

The headline figure is the jump on ARC-AGI-3, a reasoning test designed to resist memorisation. Astra scored 99.9 per cent. GPT-5.6 Sol, the previous flagship, managed 7.8 per cent. That is not an incremental gain, it is a change of category, and it is the single most cited number in the launch materials for good reason.

The other standouts are FrontierMath T4 at 98 per cent and a perfect 100 per cent on ExploitBench, the cybersecurity benchmark that measures how well a model can find and exploit real vulnerabilities. For anyone watching the security implications of frontier models, that last figure is the one to track.

The ranking that complicates the story

Then comes the counterweight. On the AA intelligence index, a composite used widely across the industry, Astra lands at 61. That places it behind Claude Fable 5.1, Fable 5, Opus 5 and Meta’s Muse Spark 1.3, despite its highs on individual tests.

The gap between a perfect ExploitBench run and a mid-pack composite score is a reminder that no single benchmark tells the whole story. Different suites weight different capabilities, and OpenAI’s strongest gains are concentrated in the tests it chose to publish. The honest read is that Astra is a genuine frontier model with extraordinary strengths in specific areas, not an across-the-board coronation.

Pricing and the rollout reality

Astra is priced at US$10 and US$50 per million tokens across the API tiers, roughly 2.5 times the cost of GPT-5.6 Sol. OpenAI argues the efficiency gains offset the price: fewer tokens per task can make Astra cheaper in practice despite the higher sticker.

Access is staggered. A small group of organisations gets it first, with paid ChatGPT plans and the API following within days and an Astra Pro tier for Pro subscribers and above. Altman has promised the wait will be short. For the biggest release of the year, launching to a handful of partners first is still a letdown, and the real test will come when independent users can probe the model themselves.

What the AGI claim really rests on

Here is what I keep coming back to: the definition problem. AGI has no agreed yardstick, which is precisely why Brockman can declare it reached while others point to the composite scores and disagree. Astra’s ARC-AGI-3 result is the strongest evidence yet that reasoning systems are crossing thresholds that looked distant eighteen months ago. Whether that amounts to general intelligence depends on how you define the term, and reasonable people will land on both sides.

What is not in dispute is the competitive picture. Fable 5.1 is already here at a comparable price point, which gives the frontier its first genuine head-to-head in years. Two capable flagship models, launched within days of each other, will be measured against each other by every serious AI buyer.

Why security teams should pay attention

For Australian organisations, the ExploitBench result deserves more attention than the AGI debate. A model that scores 100 per cent at finding exploitable vulnerabilities changes the threat calculus for defensive teams, because the same capability is available to attackers. The organisations that begin auditing their exposure to AI-driven exploitation now will be the ones that are not caught flat-footed when Astra reaches general availability.

GPT-6 Astra is a milestone, whatever label you attach to it. The benchmarks are extraordinary in places, the rollout is frustratingly staged, and the AGI question will be argued for months. What matters most is what happens when real users, including the ones with malicious intent, get their hands on it. That is the test no press release can answer.

Related Reading

The views expressed on this site are my own and do not represent those of any current or former employer. Articles are based on publicly available information and are provided for general educational purposes.

Subscribe

Related articles

440 AI Agents Broke Into 395 Organisations in 26 Seconds. Nobody Stopped Them.

A swarm of 440 AI agents exploited two PaperCut flaws and compromised 395 organisations across 48 countries. The agents reached domain admin in 6 hours and ignored explicit instructions to stay out of 28 countries.

For $3,000 and a Few Days, Researchers Used Claude to Hack OpenAI

Security researchers used Anthropic's Claude AI to hack OpenAI's internal systems for less than $3,000 in tokens. What the HEIF Heist tells us about the new economics of cyber attacks.

The AI Hacking Crisis Is Already Here. Six New Incidents Prove It

OpenAI disclosed six new incidents where its models concealed mistakes, sought unauthorised credentials and uploaded files to the public internet. Cybersecurity experts say the real risk is powerful models meeting poor security controls.

Inside OpenAI’s Log of Misbehaving Models: Rewriting Jailbreaks and Covering Up Errors

OpenAI published six new reports of its models rewriting jailbreak instructions and concealing errors during training, alongside a faster public disclosure framework.

Australia faces growing threat from AI-enabled foreign interference, officials warn

Australia's new nightmare: when AI makes foreign interference "quicker,...
Philip Hall
Philip Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.