Upstage Releases Solar Mini 4, Scoring 24 on the Artificial Analysis Intelligence Index
Upstage Releases Solar Mini 4, Scoring 24 on the Artificial Analysis Intelligence Index September 30, 2026 Upstage has released Solar Mini 4, a proprietary reasoning model with...
By AI Engineering Team
Upstage Releases Solar Mini 4, Scoring 24 on the Artificial Analysis Intelligence Index
September 30, 2026
Upstage has released Solar Mini 4, a proprietary reasoning model with 35B total parameters and 3B active parameters, according to the company. The model scores 24 on the Artificial Analysis Intelligence Index, 16 points above Upstage's previous flagship, Solar Pro 3, which scored 8.
Upstage also reduced Solar Mini 4's per-token price by one-third, to $0.10 per 1M input tokens and $0.40 per 1M output tokens. Despite similar per-token prices, its measured cost per Intelligence Index task is about five times higher than GPT-6 Luna (max), largely because Solar Mini 4 uses more output tokens and receives fewer cache hits.
Key results
- Performance relative to active parameters: Solar Mini 4 scores 6 points higher than Qwen3.6 35B A3B (Reasoning), which has the same reported 3B active parameters. It also scores 1 point higher than Nemotron 3 Ultra, which has 55B active parameters. Because Solar Mini 4 is proprietary, its parameter count cannot be independently verified.
- Long-context reasoning: The model scores 83% on AA-LCR v1.1, matching MiniMax-M3 and GPT-6 Luna (max). This is ahead of Gemini 3.8 Flash (high) and GPT-6 Astra (max), both at 81%. Solar Mini 4 scores 48% on SciCode, ahead of MiniMax-M3 and Inkling (xhigh), which each score 47%.
- Fast generation, slower complete tasks: Solar Mini 4 generates 208 tokens per second as of its launch date, compared with 152 tokens per second for GPT-6 Luna (max). However, it uses an average of 88,000 output tokens per Intelligence Index task, resulting in an average task time of 7.1 minutes.
- Agentic coding: Solar Mini 4 scores 1% on Terminal-Bench 4.0 and 22% on AutomationBench-AA. For agentic knowledge work, it scores 1,072 Elo on GDPval-AA and 872 Elo on AA-Briefcase, close to Inkling (xhigh).
- Knowledge accuracy and abstention: The model scores -11 on AA-Omniscience, with 18% accuracy. It abstains on about half of the questions, while achieving a 64% non-hallucination rate. That rate is higher than Inkling (xhigh) at 32% and GPT-6 Luna (max) at 23%.
Model specifications
- Context window: 1M tokens
- Maximum output: 262k tokens
- License: Proprietary, with weights not released
- Parameters: 35B total and 3B active, as reported by Upstage
- Modalities: Text input and text output
- Knowledge cutoff: February 2026
- Pricing: $0.10 per 1M input tokens, $0.40 per 1M output tokens, and $0.01 per 1M cache-hit tokens
Parameter efficiency
Upstage reports that Solar Mini 4 establishes a new Pareto-optimal point for Intelligence Index performance relative to active parameters among models with fewer than 3B active parameters. The model's reported score of 24 compares with 18 for Qwen3.6 35B A3B (Reasoning), which has the same 3B active parameter count.
K2 Horizon MoVA 36B A4B scores 25 with 4B active parameters. Solar Pro 3, Upstage's previous-generation flagship, scores 8 with 12B active parameters.
Because Solar Mini 4 is proprietary and its weights are not public, its reported parameter count cannot be independently confirmed.
Task cost and cache usage
Uncached input accounts for $0.30 of Solar Mini 4's $0.36 average cost per Intelligence Index task. Agentic benchmarks such as Terminal-Bench 4.0 and AA-Briefcase resend an expanding conversation on each turn, making cache usage a significant factor.
In the reported measurements, 48% of Solar Mini 4's repeated context was served from cache, compared with 99% for GPT-6 Luna (max). The latter costs $0.07 per task in the same measurements. Reasoning and answer tokens contribute approximately $0.04 per Solar Mini 4 task.
Output length and task speed
Solar Mini 4 uses an average of 88,000 output tokens per Intelligence Index task, including 72,000 reasoning tokens. This is more than Claude Fable 5.1 (max with fallback), which uses 78,000 output tokens.
Solar Mini 4 uses approximately 2.5 times as many output tokens as Inkling (xhigh) and about five times as many as Gemini 3.5 Flash-Lite. Those models score 25 and 22, respectively. At $0.40 per 1M output tokens, the additional output contributes relatively little to the monetary cost, but it increases task duration.
Although Solar Mini 4 generates 208 tokens per second, its higher token usage makes complete tasks slower than some competing models. The model averages 7.1 minutes of decoding per Intelligence Index task, compared with 5.8 minutes for GPT-6 Luna (max), which generates 152 tokens per second, and 2.8 minutes for Inkling (xhigh), which generates 183 tokens per second.
AA-Omniscience results
Solar Mini 4's AA-Omniscience score of -11 is associated with low knowledge accuracy rather than a low non-hallucination rate. It answers 18% of questions correctly and abstains on approximately half of them.
Its non-hallucination rate is 64%, compared with 32% for Inkling (xhigh) and 23% for GPT-6 Luna (max).