GPT-6 Sol and Luna Cut Inference Costs While Mixed Benchmark Results Persist
GPT 6 Sol and Luna Cut Inference Costs While Mixed Benchmark Results Persist September 22, 2026 GPT 6 Sol and GPT 6 Luna reduce the cost of running inference by roughly half com...
By AI Engineering Team
GPT-6 Sol and Luna Cut Inference Costs While Mixed Benchmark Results Persist
September 22, 2026
GPT-6 Sol and GPT-6 Luna reduce the cost of running inference by roughly half compared with GPT-5.6 Sol and GPT-5.6 Luna. Their Artificial Analysis Intelligence Index and Coding Agent Index scores remain broadly comparable with the previous generation, with improvements in some evaluations and declines in others.
Pricing and cost per task
GPT-6 Sol pricing falls from $4 per million input tokens and $20 per million output tokens to $2 and $10, respectively. GPT-6 Luna drops from $0.20 per million input tokens and $1.20 per million output tokens to $0.10 and $0.50. Both models retain the same 90% discount for cache reads and 25% premium for cache writes.
At maximum effort, GPT-6 Sol costs $1.06 per task on the Artificial Analysis Intelligence Index, approximately 50% less than GPT-5.6 Sol at $1.99. GPT-6 Luna costs $0.07 per task, about 60% less than GPT-5.6 Luna at $0.18.
The lower task cost is primarily driven by pricing. Both newer models use slightly more output tokens per task: Sol averages 31,000 tokens compared with 29,000 for GPT-5.6 Sol, while Luna averages 51,000 compared with 41,000 for GPT-5.6 Luna. The two releases therefore occupy a substantial portion of the cost-efficiency Pareto frontier.
Coding Agent Index results
In OpenAI's Codex harness, GPT-6 Sol scores 57 on the Artificial Analysis Coding Agent Index, two points higher than GPT-5.6 Sol. It improves on Terminal-Bench 4.0, scoring 43% compared with 37%, and on SWE-Atlas-QnA, scoring 58% compared with 54%.
GPT-6 Sol costs $2.99 per coding-agent task, approximately half the cost of GPT-5.6 Sol, and falls on the Pareto frontier when Coding Agent Index performance is compared with cost per task.
GPT-6 Luna scores 41, two points below GPT-5.6 Luna. Its SWE-Atlas-QnA score declines from 49% to 44%, while its DeepSWE v1.1 score falls from 66% to 64%. Despite the regression, its cost per task is approximately 60% lower.
Lower hallucination rates
Both models hallucinate less on AA-Omniscience, a benchmark measuring knowledge and hallucination behavior.
- GPT-6 Sol's hallucination rate falls from 92% to 60%.
- GPT-6 Luna's hallucination rate falls from 93% to 77%.
For Sol, the reduction is partly associated with a higher refusal rate. GPT-6 Sol attempts to answer 83% of questions, compared with 99% for GPT-5.6 Sol. This reduces wrong answers by approximately one quarter, but also lowers accuracy from 59% to 54%.
Luna's accuracy remains broadly unchanged, at 44% compared with 43%, while it answers fewer questions. On the AA-Omniscience Index, Sol rises from 22 to 27, and Luna improves from -10 to 1.
Improvements and regressions across evaluations
Outside AA-Omniscience, both models improve on AutomationBench-AA and Terminal-Bench 4.0:
- AutomationBench-AA: Sol rises from 60% to 62%, while Luna rises from 50% to 53%.
- Terminal-Bench 4.0: Sol rises from 40% to 44%, while Luna rises from 12% to 13%.
The models regress on two important knowledge-work evaluations. On GDPval-AA v2.1, a benchmark adapted from OpenAI's dataset of economically valuable tasks across 44 occupations, Sol declines by approximately 100 Elo points and Luna by approximately 75 points.
Luna also falls by roughly 45 Elo points on AA-Briefcase v1.1, while Sol remains level. AA-Briefcase v1.1 evaluates multi-week knowledge-work projects using thousands of input files.
Manual inspection of hundreds of outputs indicates that the regressions are generally associated with lower presentation quality and deliverables that omit elements required by the evaluation rubric. Both models also regress on GDPval-AA v2.1 at maximum effort, where shorter deliverables more frequently leave out required components.
Overall, GPT-6 Sol and GPT-6 Luna deliver significant reductions in cost per task and lower hallucination rates. Sol improves modestly in coding-agent performance, while Luna declines on that index. Results across broader evaluations remain mixed, with gains in automation and terminal-based tasks alongside weaker performance on selected knowledge-work benchmarks.