Skip to main content
Back to Blog
AI/MLData Analysis
2 September 20265 min readUpdated 3 September 2026

Muse Spark 1.3: Meta reaches the frontier

Muse Spark 1.3: Meta reaches the frontier Meta has released Muse Spark 1.3, its fourth Muse Spark model in five months. The release includes two variants. Muse Spark 1.3 (max),...

By AI Engineering Team

Muse Spark 1.3: Meta reaches the frontier

Meta has released Muse Spark 1.3, its fourth Muse Spark model in five months. The release includes two variants. Muse Spark 1.3 (max), currently available to Meta partners in limited preview, scores 62 on the Artificial Analysis Intelligence Index. That places it behind only Claude Fable 5.1 and Claude Opus 5. The currently available Muse Spark 1.3 (xhigh) scores 61, tying GPT-5.6 Sol (max) and Grok 4.6 (high).

The improvements in both variants are concentrated mainly in agentic work and scientific reasoning.

Intelligence Index performance

Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, four points above Muse Spark 1.2, which scored 57 in August, and eight points above Muse Spark 1.1, which scored 53 in July.

Its score ties GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high). The models ahead of it are Claude Fable 5.1 (max) at 66, Claude Opus 5 (max) at 63, and Claude Fable 5 (max) at 62.

Muse Spark 1.3 (max) reaches 62. Its higher score is supported by stronger results than the xhigh variant on Tau3-Bench Banking, where it scores 52% compared with 47%, and GDPval-AA v2, where it records 1,754 Elo compared with 1,709. The max variant is second overall only to Claude's Fable and Opus variants.

Key results

Continued gains on agentic knowledge work

Muse Spark 1.3 continues the improvement in agentic knowledge work that was visible in Muse Spark 1.2. Compared with Muse Spark 1.2, Muse Spark 1.3 (xhigh) gains:

  • 12 points on Tau3-Bench Banking, increasing from 35% to 47%
  • 5 points on Terminal-Bench 2.1, increasing from 80% to 85%
  • 94 GDPval-AA v2 Elo points, increasing from 1,615 to 1,709

Muse Spark 1.3 (max) improves further, scoring 52% on Tau3-Bench Banking and 1,754 Elo on GDPval-AA v2. Its Tau3-Bench Banking result is the highest among all models.

The max variant reaches these higher agentic-work scores by using more turns and more total reasoning tokens. Compared with the xhigh variant, it uses 62% more reasoning on GDPval-AA v2 and 28% more on Tau3-Bench Banking.

Cost per task

Muse Spark 1.3 (xhigh) costs $0.55 per Artificial Analysis Intelligence Index task under Meta's unchanged pricing of $1.25 per 1 million input tokens and $4.25 per 1 million output tokens. Cached input costs $0.15 per 1 million tokens.

Among models scoring at least 59, this is the lowest reported cost per task. GPT-5.6 Sol (max) costs $0.95 and Grok 4.6 (high) costs $0.94, more than 70% higher than Muse Spark 1.3 (xhigh). The model therefore sits on the Intelligence versus Cost per Task Pareto frontier.

Its cost per task is higher than Muse Spark 1.2, which cost $0.40 per task. The increase is driven by approximately 57% more input tokens per task on agentic evaluations, while output tokens increased by only about 8%.

Pricing for Muse Spark 1.3 (max) has not been publicly announced, so the limited-preview variant is excluded from cost comparisons.

Other models scoring 59 or higher include Gemini 3.8 Flash (high), with a score of 59 and a cost of $0.58 per task; GPT-5.6 Sol (xhigh), with a score of 59 and a cost of $0.63; and GLM-5.3 (max), with a score of 60 and a cost of $0.68. Direct peers scoring 61 cost more: Grok 4.6 (high) costs $0.94, GPT-5.6 Sol (max) costs $0.95, and Claude Opus 5 (high) costs $1.23.

Scientific reasoning

Scientific reasoning results increased across the evaluated benchmarks. CritPt produced the largest non-agentic gain for Muse Spark 1.3 (xhigh), rising eight points from 18% to 26%. GPQA Diamond increased four points, from 90% to 94%.

Humanity's Last Exam rose from 45% to 47%, while SciCode increased from 56% to 59%.

Muse Spark 1.3 (max) produced broadly similar results to the xhigh variant. It gained two points on Humanity's Last Exam, tied the xhigh variant on GPQA Diamond, and scored one point lower on CritPt.

Minor regressions

Both Muse Spark 1.3 variants declined four points on AA-LCR, from 83% for Muse Spark 1.2 to 79%.

AA-Omniscience (Accuracy) also declined. The xhigh variant fell three points, from 45% to 42%, while the max variant fell one point, from 45% to 44%. The lower scores are associated with a higher abstention rate, meaning the models more often declined to answer when uncertain. For Muse Spark 1.3 (xhigh), this also reduced the hallucination rate.

Model details: Muse Spark 1.3 (xhigh)

  • Context window: 1 million tokens, unchanged from Muse Spark 1.2
  • Pricing: $1.25 per 1 million input tokens and $4.25 per 1 million output tokens, with cached input priced at $0.15 per 1 million tokens
  • Input modalities: Text, image, and video
  • Availability: Meta's first-party API and Muse Code

Benchmark breakdown

Compared with Muse Spark 1.2, the gains for Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) on the Artificial Analysis Intelligence Index are concentrated in agentic evaluations:

EvaluationMuse Spark 1.2Muse Spark 1.3 (xhigh)Muse Spark 1.3 (max)
GDPval-AA v21,615 Elo1,709 Elo1,754 Elo
Terminal-Bench 2.180%85%86%
Tau3-Bench Banking35%47%52%

The main regressions are a four-point drop on AA-LCR for both variants, from 83% to 79%, and lower AA-Omniscience (Accuracy): three points lower for xhigh, from 45% to 42%, and one point lower for max, from 45% to 44%.

Muse Spark 1.3 (max) scores 52% on Tau3-Bench Banking, making it the highest-scoring model on that evaluation. Muse Spark 1.3 (xhigh) scores 47%, tying Claude Fable 5.1 (max) and GLM-5.3-Flash. Other higher-scoring models include Qwen3.8 Max at 51%, Grok 4.6 (high) at 51%, and GLM-5.3 (max) at 50%.

The increase from Muse Spark 1.2's 35% result is the largest individual contributor to the improved Artificial Analysis Intelligence Index scores for both Muse Spark 1.3 variants.