Skip to main content
Back to Blog
AI/MLData Analysis
3 September 20264 min readUpdated 4 September 2026

Benchmarking GPT-6 Astra

Benchmarking GPT 6 Astra GPT 6 Astra makes substantial gains in the Artificial Analysis Coding Agent Index, matching Fable 5 at less than half the cost. In the Artificial Analys...

By AI Engineering Team

Benchmarking GPT-6 Astra

GPT-6 Astra makes substantial gains in the Artificial Analysis Coding Agent Index, matching Fable 5 at less than half the cost. In the Artificial Analysis Intelligence Index, it delivers similar performance to GPT-5.6 Sol while using fewer tokens, but its higher prices outweigh those savings.

Released on September 3, 2026, GPT-6 Astra costs 2.5 times more than GPT-5.6 Sol across current pricing tiers. Input and output prices increased from $4 and $20 to $10 and $50 per million tokens. Cache reads retain a 90% discount, while cache writes carry a 25% premium.

The model produces different results across the two flagship indices. In coding-agent tasks, its token efficiency allows it to match Fable 5 at a substantially lower cost. In the Intelligence Index, improved token efficiency is offset by the price increase.

Artificial Analysis Coding Agent Index

  • Competitive with leading models: In Codex, GPT-6 Astra scores 67. This is approximately equal to Claude Opus 5 and Fable 5 in Claude Code, as well as Muse Spark 1.3 in Muse Code. Fable 5.1 in Claude Code leads the index with a score of 70.
  • About 70% more token efficient than GPT-5.6 Sol: At maximum effort in the Codex harness, GPT-6 Astra uses about one third as many tokens as GPT-5.6 Sol. It uses about one fifth as many tokens as Claude Opus 5 at the xhigh setting. Different effort levels place the model on the token-efficiency Pareto frontier.
  • Improved cost efficiency: At maximum effort, GPT-6 Astra costs approximately the same per task as GPT-5.6 Sol at maximum effort while scoring 2 points higher. It costs less than half as much per task as Fable 5 for the same score.

Coding Agent Index cost improvements are primarily associated with token efficiency. At maximum effort, GPT-6 Astra uses roughly three times fewer tokens than GPT-5.6 Sol at maximum effort.

Artificial Analysis Intelligence Index

  • Matches GPT-5.6 Sol: GPT-6 Astra scores 61, equal to GPT-5.6 Sol. This is 5 points below Claude Fable 5.1 at maximum effort with fallback. The model also trails Meta's newly released Muse Spark 1.3 at maximum effort.
  • Uses approximately 10% fewer output tokens: At maximum effort, GPT-6 Astra reduces output-token use by about 10% compared with GPT-5.6 Sol and establishes a new Pareto frontier for Intelligence Index score versus output tokens per task. However, the 2.5-fold price increase makes it 75% more expensive per task than its predecessor at maximum effort.
  • Lower hallucination rate: GPT-6 Astra records a major improvement on AA-Omniscience, the knowledge and hallucination benchmark. At maximum effort, its hallucination rate falls from 92% to 51%. Accuracy also increases by 4 points, so the reduction in hallucinations is not accompanied by lower accuracy.
  • About an 80-point gain in AA-Briefcase Elo: AA-Briefcase evaluates long-horizon knowledge work through multi-week projects involving linked tasks and thousands of source files. Compared with GPT-5.6 Sol, GPT-6 Astra improves both rubric scores and Analytical Quality Elo. Presentation Quality Elo declines, however, and GPT-5.6 Sol remains the leading model in that measure.
  • Mixed results on other evaluations: GPT-6 Astra gains 6 points on Humanity's Last Exam, which emphasizes mathematics, science, and humanities. This is offset by an approximately 80-point decline on GDPval-AA v2, adapted from OpenAI's dataset of economically valuable tasks across 44 occupations. The model also shows 2 to 3-point regressions on several other evaluations, including τ³-Banking for customer support, SciCode for scientific Python problems, and AA-LCR for long-context reasoning over large documents.

GPT-6 Astra establishes a new frontier when Intelligence Index scores are compared with output tokens per task, thanks to its roughly 10% reduction in output tokens at maximum effort. Its cost position is weaker: the model is 75% more expensive than GPT-5.6 Sol at maximum effort and generally falls behind its predecessor on the Intelligence Index versus cost-per-task frontier. The price increase drives this result, although lower token use partially offsets it.

The model's AA-Omniscience improvement is associated with the decline in hallucination rate from 92% to 51% at maximum effort, along with a modest accuracy increase.

Results on agentic knowledge work are mixed. GPT-6 Astra gains approximately 80 Elo points in AA-Briefcase but loses a similar amount in GDPval-AA v2. Within AA-Briefcase, rubric scores and Analytical Quality Elo improve, while Presentation Quality Elo declines.

The Artificial Analysis Intelligence Index v4.1.1 includes the individual evaluation results used in this comparison.