Skip to main content
Back to Blog
AI/MLData Analysis
2 September 20264 min readUpdated 22 September 2026

Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index and improves coding-agent performance

Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index and improves coding agent performance September 21, 2026 Grok 4.7 scores 46 on the Artificial Analysis Intellige...

By AI Engineering Team

Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index and improves coding-agent performance

September 21, 2026

Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index, two points higher than Grok 4.6. The evaluation used the model's xhigh reasoning effort setting and found its strongest gains in agentic knowledge work and coding-agent tasks.

Key takeaways

  • Improved agentic knowledge work: Grok 4.7 gains 111 Elo over Grok 4.6 (high) on AA-Briefcase, a private benchmark for long-horizon agentic knowledge work. It scores 1657 Elo, placing it just behind Claude Opus 5 and Claude Fable 5.1. On GDPval-AA, it scores 1695 Elo, 90 points ahead of Grok 4.6 (high).
  • Higher coding-agent score: Grok 4.7 (xhigh), used with Grok Build, scores 56 on the Artificial Analysis Coding Agent Index, nine points higher than Grok 4.6 (xhigh). Among models evaluated in their native harnesses, it ranks fourth, behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5.
  • Mixed changes on other tasks: Outside agentic knowledge work, Grok 4.7 generally performs similarly to Grok 4.6 (high). It improves by 4.5 percentage points on Terminal-Bench 4.0 and 3.0 points on GDP.pdf, while declining by 3.7 points on AA-LCR and 1.1 points on AutomationBench-AA.
  • Higher token usage: Grok 4.7 (xhigh) uses approximately 81,000 output tokens per Intelligence Index task, compared with 36,000 for Grok 4.6 (high) and 27,000 for GPT-6 Astra (max). This represents approximately 125% and 196% more tokens, respectively.

Model details

  • Context window: 500,000 tokens, unchanged from Grok 4.6
  • Pricing: $2 per 1 million input tokens and $6 per 1 million output tokens
  • Cached input pricing: $0.50 per 1 million tokens
  • Reasoning effort: Configurable from low to xhigh
  • Evaluation setting: xhigh

Agentic knowledge work performance

On AA-Briefcase, which evaluates realistic professional work tasks, Grok 4.7 scores 1657 Elo. This is an increase of 111 points over Grok 4.6 (high), placing Grok 4.7 just behind Claude Opus 5 and Claude Fable 5.1.

The improvement is concentrated in analytical quality. Grok 4.7 scores 1994 Elo for analytical quality and 1499 for presentation quality. Grok 4.6 (high) scores 1690 and 1519, respectively. AA-Briefcase also checks whether submissions satisfy task requirements, including both the analysis and the requested deliverables.

On GDPval-AA, which requires models to produce practical work products such as documents, spreadsheets, and slides, Grok 4.7 scores 1695 Elo. Grok 4.6 (high) scores 1605 Elo.

Coding-agent performance

Grok Build with Grok 4.7 (xhigh) scores 56 on the Artificial Analysis Coding Agent Index, up from 47 with Grok 4.6 (xhigh).

The model improves across all three components of the index:

  • DeepSWE v1.1: 65% to 73%
  • Terminal-Bench 4.0: 18% to 33%
  • SWE-Atlas-QnA: 58% to 63%

These results evaluate Grok with Grok Build, its first-party coding agent. They are separate from the Intelligence Index results, which use a standardized evaluation harness across models.

Token usage and evaluation time

Grok 4.7 (xhigh) scores 46 on the Intelligence Index while using approximately 81,000 output tokens per task. Grok 4.6 (xhigh) uses 38,000 tokens, while Muse Spark 1.3 (max) uses 60,000 and GPT-6 Astra (max) uses 27,000.

For long prompts, Grok 4.7's measured answer output speed is approximately 188 tokens per second. The model averages approximately 7.1 minutes per Intelligence Index task.

Hallucination and accuracy results

Grok 4.7 (xhigh) records a lower AA-Omniscience Hallucination Rate than Grok 4.6 (high), at 29% compared with 34%.

Accuracy is broadly unchanged, at 47% for Grok 4.7 and 48% for Grok 4.6. The overall AA-Omniscience Index rises from 30 to 32.

The full Intelligence Index evaluation includes results for Grok 4.7 (xhigh), Grok 4.6 (high), and other leading models.