Ant Group’s Ling-3.0-flash-Fin Scores 23 on the Intelligence Index and 24 on the Finance & Accounting Index
Ant Group’s Ling 3.0 flash Fin Scores 23 on the Intelligence Index and 24 on the Finance & Accounting Index September 16, 2026 Ant Group has released Ling 3.0 flash Fin , a fina...
By AI Engineering Team
Ant Group’s Ling-3.0-flash-Fin Scores 23 on the Intelligence Index and 24 on the Finance & Accounting Index
September 16, 2026
Ant Group has released Ling-3.0-flash-Fin, a finance-focused open weights reasoning model based on Ling-3.0-flash. The company developed the model with financial institutions and industry experts for tasks such as checking sources, creating valuation spreadsheets, and writing reports.
Ling-3.0-flash-Fin accepts and produces text only. It follows the release of Ling-3.0-flash-VL, which supports image and video input and scored 25 on the Artificial Analysis Intelligence Index.
Key results
- Ling-3.0-flash-Fin matches MiniMax-M2.7 on the Intelligence Index while using approximately half as many active parameters. Both models score 23. Flash-Fin activates 5.1B parameters per token, compared with 10B for MiniMax-M2.7.
- Ling-3.0-flash-Fin matches Ling-3.0-flash-VL on the Finance & Accounting Index, with both models scoring 24. Fin records higher business knowledge accuracy than VL, at 17% versus 11%, but also has a higher business knowledge hallucination rate, at 33% versus 19%.
- Ling-3.0-flash-Fin scores slightly below Ling-3.0-flash-VL on professional knowledge work. It scores 1171 Elo on GDPval-AA v2 and 967 on AA-Briefcase, compared with 1225 and 986 for Ling-3.0-flash-VL. Both benchmarks evaluate agents on professional tasks such as creating documents and spreadsheets.
- Difficult agentic tasks remain challenging. Ling-3.0-flash-Fin scores 7% on AutomationBench-AA, which evaluates workflows across business applications while applying guardrails. Ling-3.0-flash-VL scores 16%. Both models score 0% on Terminal-Bench v4.0, which tests difficult terminal-use tasks.
- Ling-3.0-flash-Fin produces more output tokens than both Ling-3.0-flash-VL and MiniMax-M2.7. It averages approximately 67K output tokens per Intelligence Index task, about 34% more than VL at approximately 50K and 3.2 times MiniMax-M2.7 at approximately 21K.
Model details
- Type: Open weights reasoning model
- Size: 124B total parameters, with 5.1B active per token through a mixture-of-experts architecture
- Context window: 256K tokens
- Modalities: Text input and output
- API availability: Available through OpenRouter, including a rate-limited free endpoint
- License: MIT
Parameter efficiency
Ling-3.0-flash-Fin sits on the Intelligence Index versus active parameters Pareto frontier. It scores 23 while activating 5.1B parameters per token. Ling-3.0-flash-VL scores 25 with 5.5B active parameters per token. Both models have 124B total parameters.
By comparison, Qwen3.8 27B (xhigh) scores 34 with 27B total parameters, placing both Ling models below the frontier based on total parameters.
Finance & Accounting performance
The Artificial Analysis Finance & Accounting Index combines business knowledge, reasoning, agentic work, long-context analysis, and non-hallucination. Ling-3.0-flash-Fin and Ling-3.0-flash-VL both score 24.
Fin has higher business knowledge accuracy, at 17% versus 11% for VL. However, it has a lower non-hallucination rate, at 67% compared with 81% for VL. This corresponds to business knowledge hallucination rates of 33% for Fin and 19% for VL.
Professional knowledge work
On GDPval-AA v2, which evaluates agents on professional knowledge work, Ling-3.0-flash-Fin scores 1171 Elo. This is approximately 50 points below Ling-3.0-flash-VL’s score of 1225 and above MiniMax-M2.7’s score of 1087.
Fin scores 967 Elo on AA-Briefcase, slightly below VL’s 986. AA-Briefcase tests agents on complex business workflows involving large collections of source files used to produce spreadsheets, presentations, and memos.
Compared with VL, Fin passes fewer rubric checks, 23.5% versus 24.9%, and has a lower Analytical Quality Elo score, 866 versus 907. Its Presentation Elo is slightly higher, at 1095 versus 1076, despite not having the image input capabilities of the VL model.
Token usage and evaluation results
Ling-3.0-flash-Fin averages approximately 67K output tokens per Intelligence Index task. Ling-3.0-flash-VL produces approximately 50K output tokens per task, making Fin’s average about 34% higher.
The model’s results are reported across the 10 evaluations in Artificial Analysis Intelligence Index v4.3. On AutomationBench-AA, Fin scores 7%, compared with 16% for Ling-3.0-flash-VL. Both models score 0% on Terminal-Bench v4.0.