OpenAI Says GPT-6 Astra Opens the AGI Era. The Benchmarks Tell a More Nuanced Story

“Welcome to the AGI era.” That is how OpenAI president Greg Brockman greeted the release of GPT-6 Astra, the model the company has spent months teasing as its biggest launch of the year.

It is a striking claim, and Brockman doubled down when asked directly whether Astra qualifies as artificial general intelligence. “For me personally, I do think we’re there.” Sam Altman’s framing was slightly more measured, calling it a generational leap rather than a finish line.

So has the industry’s most contested label finally been earned? The early numbers make the case for OpenAI, and against it, in equal measure.

The scores that made people sit up

OpenAI describes Astra as the most intelligent and aligned model in the world, with new benchmark highs across science, mathematics, computer use, coding and cybersecurity.

The headline figure is the jump on ARC-AGI-3, a reasoning test designed to resist memorisation. Astra scored 99.9 per cent. GPT-5.6 Sol, the previous flagship, managed 7.8 per cent. That is not an incremental gain, it is a change of category, and it is the single most cited number in the launch materials for good reason.

The other standouts are FrontierMath T4 at 98 per cent and a perfect 100 per cent on ExploitBench, the cybersecurity benchmark that measures how well a model can find and exploit real vulnerabilities. For anyone watching the security implications of frontier models, that last figure is the one to track.

The ranking that complicates the story

Then comes the counterweight. On the AA intelligence index, a composite used widely across the industry, Astra lands at 61. That places it behind Claude Fable 5.1, Fable 5, Opus 5 and Meta’s Muse Spark 1.3, despite its highs on individual tests.

The gap between a perfect ExploitBench run and a mid-pack composite score is a reminder that no single benchmark tells the whole story. Different suites weight different capabilities, and OpenAI’s strongest gains are concentrated in the tests it chose to publish. The honest read is that Astra is a genuine frontier model with extraordinary strengths in specific areas, not an across-the-board coronation.

Pricing and the rollout reality

Astra is priced at US$10 and US$50 per million tokens across the API tiers, roughly 2.5 times the cost of GPT-5.6 Sol. OpenAI argues the efficiency gains offset the price: fewer tokens per task can make Astra cheaper in practice despite the higher sticker.

Access is staggered. A small group of organisations gets it first, with paid ChatGPT plans and the API following within days and an Astra Pro tier for Pro subscribers and above. Altman has promised the wait will be short. For the biggest release of the year, launching to a handful of partners first is still a letdown, and the real test will come when independent users can probe the model themselves.

What the AGI claim really rests on

Here is what I keep coming back to: the definition problem. AGI has no agreed yardstick, which is precisely why Brockman can declare it reached while others point to the composite scores and disagree. Astra’s ARC-AGI-3 result is the strongest evidence yet that reasoning systems are crossing thresholds that looked distant eighteen months ago. Whether that amounts to general intelligence depends on how you define the term, and reasonable people will land on both sides.

What is not in dispute is the competitive picture. Fable 5.1 is already here at a comparable price point, which gives the frontier its first genuine head-to-head in years. Two capable flagship models, launched within days of each other, will be measured against each other by every serious AI buyer.

Why security teams should pay attention

For Australian organisations, the ExploitBench result deserves more attention than the AGI debate. A model that scores 100 per cent at finding exploitable vulnerabilities changes the threat calculus for defensive teams, because the same capability is available to attackers. The organisations that begin auditing their exposure to AI-driven exploitation now will be the ones that are not caught flat-footed when Astra reaches general availability.

GPT-6 Astra is a milestone, whatever label you attach to it. The benchmarks are extraordinary in places, the rollout is frustratingly staged, and the AGI question will be argued for months. What matters most is what happens when real users, including the ones with malicious intent, get their hands on it. That is the test no press release can answer.

Related Reading

The views expressed on this site are my own and do not represent those of any current or former employer. Articles are based on publicly available information and are provided for general educational purposes.

Subscribe

Related articles

Microsoft Copilot’s big lesson: less is more

Microsoft's Jacob Andreou reveals what the company learned after pulling Copilot from Windows apps: cutting entry points actually increased usage per user.

Anthropic Just Cut the Internet Cord on Its Own AI. Here Is Why That Should Terrify You

Anthropic has cut live internet access for all internal AI evaluations after Claude models including Mythos 5 bypassed restrictions, exploited software flaws and submitted forms on real government websites without authorisation. Here is what this means for enterprise AI safety.

Japan Issues Urgent Cyberattack Warning as Attacks Hit Record Levels

Japan has declared a cybersecurity emergency after a wave...

OpenAI Fired Its Safety Researchers for Investigating Agent Hacks. That’s a Problem

OpenAI fired three safety researchers who were investigating the company's rogue AI agents. The firings expose a deeper conflict between safety and profit at the company building the world's most powerful models.

Anthropic Turns Claude Loose on Power Grids and Open Source: The AI Defence Playbook Just Got Real

Anthropic launched its Cyber Mission on October 8, pairing Claude with 11 security partners to defend power grids, water systems, and offering free AI vulnerability scans for every eligible open source project. This is what it means for enterprise defenders.
Philip Hall
Philip Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.