GPT-6 Astra's AGI Benchmark Claims Don't Survive Independent Testing
OpenAI's 99.9% ARC-AGI-3 headline requires a custom harness and tens of thousands of dollars. The real score sits between 17% and 63%.
AnIntent Editorial
OpenAI wants you to read one number and stop reading: 99.9 percent on ARC-AGI-3. That figure, published alongside the September 3 launch of GPT-6 Astra, has been repeated in nearly every writeup as evidence that the AGI era has arrived. It hasn't, and the number itself doesn't mean what OpenAI is letting people assume it means.
The headline score requires a bespoke evaluation setup that ARC Prize itself does not use. Under the standard stateless harness that every other frontier model is measured against, DataCamp reports that ARC Prize's independent runs put Astra somewhere between 17 percent and 63 percent depending on reasoning tier, with the 99.9 percent figure only achievable using OpenAI's own stateful provider adapter and a comprehensive run costing tens of thousands of dollars. That is a very different claim from the one the launch materials are making.
The 99.9 Percent Number Is a Harness, Not a Capability
ARC-AGI-3 is designed to be run stateless. Each task is presented to the model cold, without memory of prior attempts, because the benchmark's purpose is to measure novel reasoning rather than accumulated task-specific inference. OpenAI's provider adapter changes that condition. It lets the model retain state across attempts, use structured tool calls the base harness doesn't expose, and burn compute at a level ARC Prize's public leaderboard has never been built to accommodate.
The result is a number that is technically true and practically misleading. Under the harness other labs use, Astra sits in a band that overlaps with strong existing models rather than transcending them. DataCamp's breakdown puts GPT-5.6 Sol at 7.8 percent and Claude Opus 5 at 30.2 percent under OpenAI's adapter conditions, but those comparisons are only meaningful if all three ran the same harness. They did not.
This is the part of the GPT-6 Astra benchmarks story that matters more than any single score: the evaluation methodology moved, and the reporting has mostly not caught up.
What Astra Actually Is
GPT-6 Astra launched September 3, 2026 as a limited preview for trusted partners, according to Wikipedia's entry tracking the release. It is a single dense reasoning model with five reasoning-effort settings and a context window in the one-million-token class, with no mini or nano tier at launch, per ComputingForGeeks. The knowledge cutoff on the model card is April 30, 2026, though the model itself misreports this date.
Pricing is where the strategic message gets clearer. Yotta Labs' summary lists API rates at $10 per million input tokens and $50 per million output, with cached input at $1 and batch runs at half price, plus a Fast mode at 2x rate. That works out to roughly 2.5x GPT-5.6 Sol's promotional pricing. ComputingForGeeks characterizes the pricing as brutal and explicitly not a drop-in upgrade.
A 2.5x price hike is not what a lab does when it believes it has a decisive capability lead across every task. It is what a lab does when it has a decisive lead on some tasks and needs the pricing to signal tier separation.
The Reasoning Trace Went Dark, and That Matters More Than the Score
Here is the detail most launch coverage skipped: Astra uses a new technique called recurrent depth, or looped transformers, that Wikipedia documents as obscuring some or all of the model's chain of thought, quoting Russell Brandom's reporting in The Verge on September 2, 2026. ComputingForGeeks describes the reasoning as harder to inspect than any model before it, with gains that are real but uneven.
Inspectable reasoning has been the field's primary safety and debugging tool for the last two years. Interpretability research, red-teaming methodology, and enterprise auditing workflows all assume that a reasoning model produces a legible trace of how it arrived at its answer. Recurrent depth breaks that assumption by design. The model reasons in latent space through repeated passes over its own internal states, and the token-level chain of thought that used to explain the answer is no longer a faithful record of the computation.
A lab claiming AGI-era capability while simultaneously shipping an architecture that makes those capabilities harder to audit is asking the industry to accept both claims on faith. Neither claim is verified. That is the tension no benchmark chart resolves.
Where the GPT-6 Astra vs Claude Opus 5 Comparison Actually Lands
On the benchmarks where methodology is less contested, Astra does look like a real step forward. DataCamp lists a 97.6 percent score on FrontierMath Tier 4 and 72.6 percent on OSWorld 2.0, with roughly 47 percent less time per task than GPT-5.6 Sol, both vendor-reported. On ExploitBench, running June through August 2026 with contamination controls, Astra scores 39.0 percent against Sol's 5.5 percent, and on SRE-Bench it solves 88.0 percent in one attempt versus 55.9 percent for Sol.
The cybersecurity numbers are the ones that should be getting the attention the ARC-AGI-3 figure is getting. Independent lab Irregular, cited in DataCamp's report, had Astra solving 86 of 226 FrontierCyber challenges compared with 34 for Sol, including zero-day findings in browsers and a cloud database. That is a categorically different capability profile from prior models, and it is why OpenAI reported on September 1, 2026 that access to its most advanced cybersecurity capabilities would be more limited, with advanced cyber access initially restricted to vetted testers through the Daybreak Blue program, per Wikipedia.
Claude Fable 5.1, released September 1, 2026, and Claude Opus 5 remain the reference points for coding and agentic work at the frontier, DataCamp notes. Nothing in the launch materials suggests Astra has displaced them across those workloads. It has opened a lead in reasoning-heavy tasks and cyber offense, and it has raised the price to match.
The Best Version of the Opposing View, and Why It Still Falls Short
The strongest defense of the AGI-era framing goes like this: a model that can score 99.9 percent under any legitimate harness has demonstrated a capability that did not previously exist, regardless of cost or configuration. The 17 to 63 percent stateless score isn't the ceiling of what Astra can do, it's the floor of what Astra can do without help. If the model can reach near-perfect performance with the right scaffolding, then the underlying reasoning capability is present, and the harness question is an implementation detail.
That argument would carry weight if the scaffolding were reproducible by third parties at reasonable cost. It isn't. A comprehensive run under the adapter configuration costs tens of thousands of dollars, per DataCamp, and the adapter itself is OpenAI's own tooling rather than a standardized evaluation environment. Capability that only manifests under a vendor-controlled, five-figure-per-run configuration is a capability the market cannot independently verify. It also cannot be priced, deployed, or safety-tested by anyone outside the vendor's evaluation loop.
The historical parallel is worth naming. GPT-4's original MMLU numbers looked one way in the technical report and another way when independent labs re-ran them under matched conditions. The gap closed, but only because the benchmark methodology was public and cheap enough to reproduce. ARC-AGI-3 under OpenAI's adapter is neither.
OpenAI's AGI Era Claims Need a Different Kind of Evidence
The OpenAI AGI era claims that have accompanied Astra's rollout are doing two jobs at once: they are marketing to enterprise buyers weighing a 2.5x price increase, and they are political signaling to regulators about the seriousness of frontier capability. Both jobs are easier when the top-line benchmark reads 99.9 percent. Neither job is honest without the harness footnote.
BenchLM currently ranks Astra at 81.1 out of 100, second of 232 models as of September 4, 2026, with reasoning as its strongest category at number one. Independent runtime speed has not yet been measured. That is a serious model. It is not a model that has ended the benchmark era, and the fact that independent runtime numbers don't exist yet is itself a signal about how new and how partially-observed this release is.
The rollout structure reinforces the point. Emergent.sh's release tracking notes that a named model, a research post, and a live enterprise preview all landed at once, but broad public access did not. It is a phased release, not an open launch. Yotta Labs confirms enterprise access is off by default until an admin enables it, with ChatGPT Plus, Pro, Business, and Enterprise scheduled over the coming days, plus OpenAI API and AWS Bedrock. For a model being described in terms that imply general intelligence, the deployment posture is unusually cautious. Correctly so, given the cyber capability profile.
Those interested in how the safety framing shapes deployment can read our coverage in AI Safety articles and the enterprise-side implications in Enterprise AI articles. For the cyber angle specifically, AI Cybersecurity articles tracks the pattern of labs restricting offensive capability tiers.
What to Watch Next
The useful signal over the next 60 days will not be new marketing benchmarks. It will be whether ARC Prize publishes its full independent run of Astra under the standard stateless harness, with reasoning tier disclosed and cost transparent. If that number lands near the top of the 17 to 63 percent range and Astra keeps its lead on FrontierMath and ExploitBench under contamination controls, the model deserves the tier-one framing without the adapter footnote. If it lands near the bottom, the 99.9 percent headline becomes the defining example of harness-specific benchmarking in the frontier era.
Either way, treat the GPT-6 Astra ARC-AGI-3 score with the caveat OpenAI's own methodology requires: it is a number produced under a configuration the rest of the industry cannot replicate at comparable cost. Ask your vendor which harness their evaluation used before you agree to the 2.5x price. That question, not the score, is what separates buyers who understand what they are getting from buyers who are paying for a chart.
Frequently Asked Questions
What is the difference between a stateless and stateful harness in AI benchmarks?
A stateless harness presents each task cold with no memory of prior attempts, which is how ARC-AGI-3 is designed to be run. OpenAI's provider adapter for Astra lets the model retain state across attempts and use structured tool calls, which is why its 99.9% score is not directly comparable to other models' stateless runs.
How much does it cost to run GPT-6 Astra via the API?
Yotta Labs lists API pricing at $10 per million input tokens and $50 per million output tokens, with cached input at $1 and batch runs at half price. A Fast mode is available at 2x the standard rate, and the overall pricing is roughly 2.5x GPT-5.6 Sol's promotional rate.
Why is OpenAI restricting Astra's cybersecurity capabilities?
OpenAI classifies Astra as meeting the Critical cybersecurity threshold under its Preparedness Framework, per Yotta Labs. Advanced cyber access is initially restricted to vetted testers through the Daybreak Blue program because Irregular's independent testing found Astra capable of zero-day discovery in browsers and a cloud database.
What is recurrent depth and why does it matter for AI safety?
Recurrent depth, also called looped transformers, is the new reasoning technique Astra uses. According to Wikipedia citing The Verge, it obscures some or all of the model's chain of thought, meaning the token-level reasoning trace that safety researchers and enterprise auditors rely on is no longer a faithful record of how the model computed its answer.
When will GPT-6 Astra be available to regular ChatGPT users?
The rollout started September 3, 2026 with limited organizations and enterprise customers in the Daybreak program, per Emergent.sh. Yotta Labs indicates ChatGPT Plus, Pro, Business, and Enterprise access will follow over the coming days, along with availability on the OpenAI API and AWS Bedrock, with enterprise access off by default until an admin enables it.
Written by
AnIntent Editorial
AnIntent is an independent technology and automotive publication. Our editorial team researches every article from live primary sources, cross-checks key facts across multiple references, and cites claims inline so readers can verify them directly. We cover smartphones, laptops, EVs, gaming hardware, AI tools, and more — with no sponsored content and no paid placements.