Skip to main content
Back to Blog
AI/MLInnovation
1 October 20264 min readUpdated 1 October 2026

Gemini 4 Argon Places Google Among the Top Three AI Labs by Intelligence Score

Gemini 4 Argon Places Google Among the Top Three AI Labs by Intelligence Score September 30, 2026 Google DeepMind's Gemini 4 Argon is the company's first proprietary model above...

By AI Engineering Team

Gemini 4 Argon Places Google Among the Top Three AI Labs by Intelligence Score

September 30, 2026

Google DeepMind's Gemini 4 Argon is the company's first proprietary model above the Flash class in more than seven months. With high reasoning enabled, it scored 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and placing one point ahead of GPT-6.1 Sol (max, 52).

The result represents a 23-point improvement over Google's previous non-Flash model, Gemini 3.1 Pro Preview, which scored 30. Gemini 4 Argon also scored 12 points higher than Gemini 3.8 Flash (high).

Gemini 4 Argon is being rolled out to selected users and is not yet publicly available. Its current pricing includes a 50% introductory discount, although Google has not confirmed when the promotion will end.

Cost per task

At the discounted rates, Gemini 4 Argon costs $1.99 per Intelligence Index task. That is 60% of GPT-6 Astra (max), which costs $3.26 per task, for a comparable Intelligence Index score.

The cost advantage comes from lower token prices rather than lower token usage. Gemini 4 Argon averages 62,000 output tokens per task, compared with 27,000 for GPT-6 Astra (max).

At standard pricing, Gemini 4 Argon's cost per task is expected to rise to $3.98, approximately 1.2 times the cost of GPT-6 Astra (max). The model costs $4 per 1 million input tokens and $20 per 1 million output tokens at standard rates. During the introductory discount, those prices are reduced to $2 and $10, respectively.

Cached input tokens receive a 95% discount, costing $0.10 per 1 million tokens at the discounted rate. This is an increase from the 90% cache discount available for Gemini 3.8 Flash.

Agentic benchmark performance

Gemini 4 Argon shows improved results across agentic evaluations, an area where earlier Gemini models had generally been weaker.

  • On AutomationBench-AA, it ranks first with a score of 78%, seven points ahead of Claude Sonnet 5.5 (max), which scored 71%.
  • On Terminal Bench 4, it scored 57%, a 53-point improvement over Gemini 3.1 Pro Preview. It ranked behind Claude Sonnet 5.5 (max, 64%), Claude Opus 5.5 (max, 60%), and GPT-6 Astra (59%).
  • On AA-Briefcase, it achieved 1494 Elo.

The AA-Briefcase result includes a 65% rubric pass rate, the highest recorded in the evaluation. Its Analytical Quality score was 1576 Elo, while its Presentation Quality score was 1308 Elo.

Hallucination and accuracy results

On AA-Omniscience, Gemini 4 Argon recorded a 15% hallucination rate. This was the lowest rate among models scoring at least 45 on the Intelligence Index. GPT-6 Astra (max) recorded 51%, while GPT-6.1 Sol (max) recorded 54%.

The model's accuracy score was 50%, five points lower than Gemini 3.1 Pro Preview and 13 points below GPT-6 Astra (max), which scored 63%. Despite the lower accuracy score, Gemini 4 Argon's overall AA-Omniscience score was 42, comparable to GPT-6 Astra's 43 and GPT-6.1 Sol's 42.

The lower hallucination rate indicates that Argon is more likely to acknowledge uncertainty instead of producing an incorrect guess. Accuracy and hallucination rate are measured separately in the evaluation.

Model details

  • Context window: 1 million tokens
  • Input modalities: Text, image, video, and speech
  • Output modality: Text
  • Standard pricing: $4 per 1 million input tokens and $20 per 1 million output tokens
  • Introductory pricing: $2 per 1 million input tokens and $10 per 1 million output tokens, with the discount available for at least one month
  • Cached input pricing: $0.10 per 1 million tokens at the discounted rate

Long Decode Continuation

Gemini 4 Argon was also tested with Long Decode Continuation, a Gemini API feature that pauses long responses and resumes them through follow-up calls. The feature allows reasoning to continue for up to 1 million output tokens without request timeouts.