Skip to main content
Back to Blog
AI/MLData Analysis
1 September 20264 min readUpdated 2 September 2026

Google Releases Gemini 3.8 Flash, Its Fourth Flash Model in Under Four Months

Google Releases Gemini 3.8 Flash, Its Fourth Flash Model in Under Four Months Google DeepMind released Gemini 3.8 Flash on September 2, 2026. At its high reasoning setting, the...

By AI Engineering Team

Google Releases Gemini 3.8 Flash, Its Fourth Flash Model in Under Four Months

Google DeepMind released Gemini 3.8 Flash on September 2, 2026. At its high reasoning setting, the model scores 59 on the Artificial Analysis Intelligence Index, a three-point increase over Gemini 3.7 Flash’s score of 56.

The score matches the sub-maximum reasoning results of GPT-5.6 Sol at xhigh and Grok 4.6 at medium, both of which also score 59.

Gemini 3.8 Flash uses Gemini 3.7 Flash’s discounted pricing through the end of 2026: $0.75 per million input tokens and $3.75 per million output tokens. At high reasoning, its average cost per task is $0.58, placing it on the Intelligence vs. Cost per Task Pareto frontier. That cost is comparable to GPT-5.6 Terra at maximum reasoning, which costs $0.53 per task.

However, Gemini 3.8 Flash’s cost per task is approximately 40% higher than Gemini 3.7 Flash’s $0.40. The increase is associated with a 30% rise in average output tokens per task, reaching 48,000 tokens, as well as more turns in agentic evaluations.

Benchmark results across reasoning levels

Gemini 3.8 Flash is available with high, medium, and low reasoning levels. The main benchmark results are:

  • High reasoning: The model scores 59 on the Artificial Analysis Intelligence Index, three points above Gemini 3.7 Flash’s high-reasoning score of 56.
  • Medium reasoning: It scores 57, matching GPT-5.6 Terra at maximum reasoning and Muse Spark 1.2 at xhigh reasoning.
  • Low reasoning: It scores 52, matching Gemini 3.6 Flash at high reasoning. At this setting, Gemini 3.8 Flash has a 30% lower cost per task and approximately one-third of the time per task.

The three-point improvement is attributed primarily to stronger results on agentic evaluations, including 𝜏³-Banking for tool use, Terminal-Bench v2.1 for coding, and GDPval-AA v2 for real-world tasks.

The largest gain appears on 𝜏³-Banking, where Gemini 3.8 Flash improves by 12 points over Gemini 3.7 Flash and reaches 45%.

Cost and task efficiency

At high reasoning, Gemini 3.8 Flash costs $0.58 per Intelligence Index task, making it the least expensive model at its level of intelligence in the reported comparison. Its cost is approximately 40% higher than Gemini 3.7 Flash’s $0.40, despite unchanged per-token pricing.

The increase reflects higher token usage and additional turns on agentic evaluations. Cost per task falls to $0.41 with medium reasoning and $0.24 with low reasoning.

Speed and time per task

Gemini 3.8 Flash remains fast in terms of output generation. At high reasoning, it averages approximately 300 output tokens per second and has a time per task of 2.5 minutes.

That is slightly faster than GPT-5.6 Luna at maximum reasoning, which takes 2.6 minutes, and GPT-5.6 Terra at maximum reasoning, which takes 3.3 minutes. Gemini 3.8 Flash is slower than Claude Fable 5.1 at medium reasoning, which takes 2.1 minutes.

Compared with Gemini 3.7 Flash, the higher token usage raises time per task from 2.2 minutes to 2.5 minutes. At low reasoning, time per task drops to 0.8 minutes, placing Gemini 3.8 Flash on the Intelligence vs. Time per Task Pareto frontier.

Model specifications and pricing

  • Context window: 1 million tokens, unchanged from Gemini 3.7 Flash
  • Input modalities: Text, image, video, and speech
  • Output modality: Text
  • Discounted pricing through the end of 2026: $0.75 per 1 million input tokens and $3.75 per 1 million output tokens
  • Standard pricing: $1.50 per 1 million input tokens and $7.50 per 1 million output tokens
  • Cached input tokens: A 90% discount remains available