Skip to main content
Back to Blog
AI/MLInnovation
9 August 20263 min readUpdated 28 August 2026

Agnes 2.5 Pro Beta reaches 49 on the Artificial Analysis Intelligence Index

Agnes 2.5 Pro Beta reaches 49 on the Artificial Analysis Intelligence Index August 27, 2026 Agnes AI's Agnes 2.5 Pro Beta scored 49 on the Artificial Analysis Intelligence Index...

By AI Engineering Team

Agnes 2.5 Pro Beta reaches 49 on the Artificial Analysis Intelligence Index

August 27, 2026

Agnes AI's Agnes 2.5 Pro Beta scored 49 on the Artificial Analysis Intelligence Index, improving by 9 points from Agnes 2.5 Pro Alpha's score of 40. The increase was driven primarily by stronger agentic performance, although the beta model used approximately twice as many output tokens per benchmark task.

Agnes AI is a Singapore-based AI lab that develops full-modality foundation models and provides them through a free omni-modal API. Agnes 2.5 Pro Beta is a beta checkpoint for the next 2.5 Pro release. It accepts text and image inputs and produces text output.

With a score of 49, the model moved from the middle of the rankings into a tier close to the frontier. It ranked just below Gemini 3.5 Flash (high, 52) and GPT-5.6 Luna (max, 52), while scoring above MiniMax-M3 (45).

Key results

  • Intelligence Index: Agnes 2.5 Pro Beta scored 49, up from 40 for Agnes 2.5 Pro Alpha.
  • Agentic Index: The score increased from 25 to 44. This placed the model just behind Gemini 3.7 Flash (high, 45), and ahead of Gemini 3.5 Flash (high, 40) and MiniMax-M3 (36).
  • τ³-Banking: Performance nearly tripled, rising from 12% to 36%.
  • GDPval-AA v2: The model's Elo increased from 1171 to 1456, against a human baseline of 1000.
  • Humanity's Last Exam: The score rose from 34% to 38%.
  • GPQA Diamond: The score improved from 88% to 91%.
  • CritPt: The score increased from 11% to 16%.

The largest improvements came from agentic evaluations. On GDPval-AA v2, Agnes 2.5 Pro Beta reached an Elo of 1456, compared with 1171 for Agnes 2.5 Pro Alpha. This result was ahead of MiniMax-M3 at 1384 and behind GPT-5.5 (xhigh, 1489) and Gemini 3.7 Flash (high, 1527).

Abstention affected the AA-Omniscience result

Agnes 2.5 Pro Beta's AA-Omniscience score improved from -25 to -11, but the change reflected greater abstention rather than higher answer accuracy.

The beta model attempted 45% of the questions, compared with 94% for Agnes 2.5 Pro Alpha. Its hallucination rate fell from 88% to 33%, while its AA-Omniscience Accuracy score was reduced from 33% to 17%.

Token usage

The improvement in the Intelligence Index score required substantially more output. Agnes 2.5 Pro Beta used approximately 50,000 output tokens per Intelligence Index task, more than twice Agnes 2.5 Pro Alpha's 24,000 tokens.

The beta model also used more tokens than Qwen3.8 27B (xhigh, 47,000) and GLM-5.3 (max, 41,000).

Model details

  • Context window: 1M tokens
  • Maximum output: 65k tokens
  • Input modalities: Text and image
  • Pricing: $0.10 per 1M input tokens, $0.30 per 1M output tokens, and $0.01 per 1M cache-hit tokens
  • Availability: Agnes AI first-party API

Overall, Agnes 2.5 Pro Beta's 9-point Intelligence Index improvement was associated mainly with gains in agentic tasks. Its performance on frontier reasoning evaluations increased more modestly, while the model's higher token usage and AA-Omniscience abstention rate remain important factors in interpreting the results.