AI · · 4 min read
GPT-6 Astra Sets New ARC-AGI-3 Benchmark Record
OpenAI’s GPT-6 Astra leads ARC-AGI-3 by a wide margin, though its results vary sharply depending on how the evaluation preserves reasoning state.
OpenAI’s GPT-6 Astra has set a new high on the ARC-AGI-3 reasoning benchmark, scoring 62.7% under the neutral testing system used by ARC Prize. The result is more than twice the previous verified frontier score, according to reporting by Startup Fortune.
The achievement came at considerable expense. A complete evaluation at Astra’s highest reasoning setting cost about $26,000, based on the estimate supplied with the test results. The score puts Astra well ahead of Claude Opus 5, which reached 30.2% in July, and OpenAI’s previous leading model, GPT-5.6 Sol, which recorded 7.8%.
The result is also more meaningful because it was produced under the Standard harness, which is designed to apply the same basic conditions to models from different companies. That makes Astra’s score of 62.7% the most useful figure for comparing it with competitors, even though OpenAI obtained a much higher result under its own testing system.
Two tests, two very different results
When Astra was evaluated through OpenAI’s Provider Adapter harness, its best observed result rose to 99.9%. That run used the model at high reasoning and cost approximately $18,800. The change was not a new version of Astra, but a change in the way the model’s internal reasoning was carried from one step to the next.
ARC-AGI-3 presents models with interactive puzzles that require them to understand unfamiliar environments and act within them. In the Standard harness, a model can preserve notes that it deliberately creates, but its private reasoning state is not retained between requests. Every model on the public leaderboard faces that restriction.
OpenAI’s adapter allows Astra to retain opaque reasoning state throughout a puzzle and to compress that state when necessary. That additional continuity produced the 37.2-point difference between the two results. ARC Prize has identified the distinction in its analysis, meaning the 99.9% figure should not be treated as directly equivalent to the neutral leaderboard score.
For cross-company comparisons, the Standard result therefore remains the more defensible benchmark. Even on those stricter terms, Astra’s lead is substantial. Anthropic has not released a verified ARC-AGI-3 result for Claude Opus 5.5. Grok 4.6 scored 2.11% in an XHigh run, while earlier Grok 4.5 entries reached 0.3%; no verified Grok 4.7 result appeared on the leaderboard at the time covered by the report.
Evidence of more efficient problem-solving
The ARC Prize analysis looked beyond the percentage score. On 96% of the levels tested, Astra used fewer actions than the median human participant. Across those tasks, it averaged 51.7% fewer moves before reaching a solution.
The model also appeared to construct compact symbolic descriptions of unfamiliar game worlds and reuse them when planning subsequent actions. That approach differs from repeatedly probing an environment without retaining a coherent working model. It more closely resembles learning the rules of a new game and applying that understanding as play continues.
Those findings do not establish exactly how Astra is built. OpenAI has not provided a detailed account of the model’s architecture. Aidan Clark, the company’s vice-president of research, said Astra was the first OpenAI model trained in a run involving more than 100,000 GPUs at the company’s Stargate site in Texas. He also said it was the first training run in which earlier OpenAI models helped supervise the development of their successor.
AI researcher Sebastian Raschka has discussed reports that Astra may use recurrent depth, sometimes described as looped transformers, but OpenAI has not confirmed that design. The company has said Astra handles more difficult tasks while verbalising less of its reasoning and gives tighter control over the reasoning trace shown to users.
Performance gains and reduced visibility
That trade-off raises questions about oversight. If more of a model’s reasoning remains hidden, outside observers have less material with which to examine how it reached an answer. OpenAI’s own information acknowledges that Astra is harder to monitor through visible chain-of-thought inspection than earlier models, despite its stronger performance.
Astra was announced on September 3 and subsequently made available through ChatGPT Plus, Pro, Business and Enterprise plans, as well as the API, Microsoft Azure and AWS Bedrock. Its listed pricing is $10 per million input tokens and $50 per million output tokens, with a context window of 1.05 million tokens.
The ARC-AGI-3 result adds a new measure of Astra’s capabilities after launch. It suggests that the model’s advantage is especially pronounced on tasks requiring an agent to form and maintain a model of an unfamiliar environment. At the same time, the large difference between neutral and provider-specific testing shows how strongly benchmark outcomes can depend on the tools and memory available during evaluation.
For now, the comparable scoreboard places Astra at 62.7%, Claude Opus 5 at 30.2% and the latest verified Grok result at 2.11%. The next question is whether Anthropic and xAI can close that gap under the same testing conditions.