Ling 3.1 Flash, InclusionAI's September 2026 release, scored 41 on the Artificial Analysis Intelligence Index, roughly double its predecessor's 20. It suits teams prototyping coding and tool-using agents on a free trial, but anyone needing downloadable weights, vision input or published privacy terms should wait or choose a rival.
Ling 3.1 Flash is a 560-billion-parameter mixture-of-experts reasoning model from InclusionAI, Ant Group's AI lab, that activates about 25 billion parameters per token. It accepts and returns text only, and its trial hosting serves a 262,144-token context while the weights remain unreleased.
Where it sits
- $0.45/M$ per 1M tokensBlended price (3:1)Lower is better#19 / 79peer median $1.69/Mvendor price, checked by HokAI
- 211 tok/stokens/sOutput speedHigher is better#13 / 50peer median 90 tok/scited: Artificial Analysis
Cheaper than 77% of the 79 GA models with a published price, and rank 13 of 50 on output speed as cited from Artificial Analysis. Ranked against GA models; this record is not GA.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: InclusionAI · Family: Ling 3.1
More about InclusionAI on HokAI
Context window: 262,144 tokens · Max output: 32,768
Input modalities: text, tool-calls · Output: text, tool-calls
About Ling 3.1 Flash
InclusionAI, the AI lab inside Ant Group (see InclusionAI), announced Ling-3.1-flash on 30 September 2026, China time. It follows Ling 3.0 Flash from July, which had 124 billion total parameters and about 5.1 billion active, and it sits on a different branch from the finance fine-tune Ling 3.0 Flash Fin. The new model is a sparse mixture of experts with roughly 560 billion total and 25 billion active parameters. InclusionAI positions it for agent tasks, search, office software and specialist applications. The lab's other releases are collected on the InclusionAI provider page.
The independent evidence comes from Artificial Analysis, which ran its ten-evaluation Intelligence Index and scored the model 41. Inside that index it reached 33.3 percent on Terminal-Bench 4.0, 39.4 percent on Humanity's Last Exam and 54.1 percent on SciCode, plus an Elo of 1,622 on GDPval-AA, its test of real knowledge-work deliverables. On AA-Omniscience it answered 29.1 percent of questions correctly and recorded a 37.9 percent hallucination rate, Artificial Analysis's measure of how often a model gives a wrong answer instead of abstaining when it does not know. Third-party write-ups quote higher vendor figures taken from launch graphics we could not inspect, so every score on this page comes from the Artificial Analysis run.
On InclusionAI's own API, Artificial Analysis measured a median 1.79 seconds to the first chunk and about 13.6 seconds end to end on its long-prompt test, of which roughly 9.5 seconds was reasoning before the answer began. The index run consumed about 218 million output tokens, more than twice the 100 million median Artificial Analysis cites for comparable models, and cost about 99 cents per task at the listed rate.
The hosted trial endpoints serve a 256K-token window and up to 32,768 output tokens. InclusionAI says the full 1M-token window arrives when the trial ends and the service turns paid. Input and output are text only, which places it among the text-only models in the directory. Novita AI and OpenRouter expose function calling with tool_choice and a reasoning parameter that is on by default, and Novita adds an Anthropic-style messages endpoint beside chat completions.
On 8 October 2026 the model was free to call on Novita AI, OpenRouter and Vercel AI Gateway. Those are promotional rates, not the list price, so the trial is the place to test before billing starts. The same hosts appear in the AI gateways category, and the AI agent builders category covers the frameworks that would sit in front of an agent model like this one. Compare it with GLM-5.3 Flash, DeepSeek V4.1 Flash, Gemini 3.8 Flash and Claude Haiku 5.5 before committing. On HokAI's own records GLM-5.3 Flash lists a lower rate, a slightly higher index score and MIT-licensed weights; DeepSeek V4.1 Flash lists a lower rate with MIT weights and a slightly lower score; Gemini 3.8 Flash scores about the same, costs more and takes image, audio and video input; Claude Haiku 5.5 lists a lower rate and a higher score. Further up the index, DeepSeek V4 Flash lists a lower rate and a higher score with open weights, Tencent's Hy3 scores far higher at a similarly low rate, and Kimi K3 scores higher at a much higher rate. The model leaderboard sorts all of them on the same index.
InclusionAI says it plans to open-source the model after the trial but has not named a license, a date or a checkpoint format. On 8 October 2026 the expected Hugging Face page at inclusionAI/Ling-3.1-flash returned a 404, Artificial Analysis classes the model as proprietary, HokAI files it as a preview release while the trial runs, its hosted context sits in the 128K to 400K bucket, and we found no system card, data-retention statement or training-data summary. Ling 3.0 Flash and its Fin fine-tune both have public Hugging Face repositories, and the Fin model carries the MIT license, which is a pattern, not a promise.
The sensible use today is a time-boxed trial: run your own terminal and tool-calling tasks while the hosts are free, cap output length because the model is verbose, and decide on the list rate afterwards. Teams that need to self-host, send images or sign a data-processing agreement should look at the peers above until InclusionAI publishes weights and terms.
Screenshots

Pricing
Artificial Analysis recorded a rate of $0.30 per million input tokens and $0.90 per million output tokens, with cached input at $0.06 (an 80% discount). The launch trial is free: Novita AI, OpenRouter and Vercel AI Gateway all showed $0 on 8 October 2026, and InclusionAI described the free period as two weeks, after which the service turns paid. At the listed rate a request of one million input and one million output tokens costs $1.20. [GLM-5.3 Flash](/hub/models/glm-5.3-flash) and [DeepSeek V4.1 Flash](/hub/models/deepseek-v4.1-flash) list lower rates on their HokAI records, so the case for Ling is speed and agent behavior, not price. The model is verbose, which pushes the real bill above a per-token comparison.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.0090 | $0.0009 | $0.0099 |
| Support reply | $0.0006 | $0.0003 | $0.0009 |
| One coding agent run | $0.060 | $0.018 | $0.078 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- Sparse mixture of experts: Only about 4.5% of the weights run for each token, the design choice behind its measured speed.
- Fast output: Artificial Analysis measured a median 211.4 output tokens per second on InclusionAI's own API.
- Reasoning on by default, switchable: Thinking runs unless you turn it off with the reasoning parameter; Artificial Analysis saw about 9.5 seconds of reasoning before the first answer token on its long-prompt test.
- Terminal and tool work: Function calling with tool_choice works on Novita AI and OpenRouter, and the Terminal-Bench 4.0 run scored 33.3%.
- Long-context reasoning: It scored 83% on AA-LCR, Artificial Analysis's long-context reasoning test.
- Anthropic-style endpoint: Novita lists an Anthropic messages endpoint beside chat completions, so Claude-format clients can point at it.
Pros
- Artificial Analysis ran its full ten-evaluation index within days of launch, so the headline score is measured by an outside party, not taken from vendor marketing.
- Speed barely changes between medium and long prompts in Artificial Analysis's tests, which suits long agent loops.
- Free to call on three hosts during the trial, so evaluating it costs only integration time.
- InclusionAI says weights will follow the trial, and its earlier finance fine-tune was published under MIT.
Cons
- Verbose: Artificial Analysis counted about 218 million output tokens on its index run, more than double the 100 million median it cites, which raises real spend on a per-token plan.
- Weights, license, system card and data policy are all unpublished, and no Hugging Face repository existed for it on 8 October 2026.
- Text only: it cannot take images, audio clips or video.
- Hosted trials cap context at 262,144 tokens, a quarter of the 1M design target.
Benchmarks
- GDPval-AA v2: 1,622 cited: Artificial Analysis · 08 Oct 2026 — Real knowledge-work deliverables judged against professionals, run by Artificial Analysis.
- Humanity's Last Exam: 39.4% cited: Artificial Analysis · 08 Oct 2026 — Expert-written questions across many fields, % correct.
- AA Intelligence Index: 41.1 cited: Artificial Analysis · 08 Oct 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
- AA blended price: $0.45/M cited: Artificial Analysis · 08 Oct 2026 — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.
- Output speed: 211 tok/s cited: Artificial Analysis · 08 Oct 2026 — Median tokens written per second as measured by Artificial Analysis.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
How much do you pay for Ling 3.1 Flash?
Artificial Analysis records $0.30 per million input tokens, $0.90 per million output tokens and $0.06 for cached input, measured on InclusionAI's own API. During the launch trial Novita AI, OpenRouter and Vercel AI Gateway charge nothing, so a test run cost $0 on 8 October 2026. Against Claude Haiku 5.5's HokAI-listed $0.10 and $0.50, the list rate is three times higher on input and almost twice as high on output.
Is Ling 3.1 Flash better than GLM-5.3 Flash?
On the numbers HokAI holds, no: GLM-5.3 Flash scores slightly higher on the Artificial Analysis index, lists a lower per-token rate, accepts images, video and PDFs, and ships MIT-licensed weights you can download. Ling's case rests on its measured output speed, its agent-oriented tuning and the free trial. Pick GLM unless you have a specific reason to test Ling.
Is Ling 3.1 Flash open source?
Not yet. InclusionAI says it plans to release the weights after the free trial, but no license, date or checkpoint format has been published, and the inclusionAI/Ling-3.1-flash page on Hugging Face returned a 404 on 8 October 2026. Artificial Analysis currently labels it proprietary, so treat it as an API-only model until a repository appears.
Does InclusionAI train on what you send to Ling 3.1 Flash?
InclusionAI has not published a retention or training policy for this model that we could find. When you call it through Novita AI, OpenRouter or Vercel AI Gateway, that host's own terms apply to your prompts, so read them before sending customer or proprietary data. Until a policy exists, keep regulated or confidential material out of the trial.
Who should try Ling 3.1 Flash right now?
Developers building terminal or tool-calling agents who can cap output length and want a no-cost test of a large reasoning model. Pass on it if you need vision input, a self-hosted checkpoint or a signed data agreement, and look at GLM-5.3 Flash or Gemini 3.8 Flash instead. Its verbosity also makes it a poor fit for long-form chat where token spend is the main cost.
Top Alternatives
- GLM-5.3 Flash: Pick GLM-5.3 Flash if you want published MIT weights and a lower per-token rate today; pick Ling 3.1 Flash to test InclusionAI's agent-focused tuning on a free trial.
- DeepSeek V4.1 Flash: Pick DeepSeek V4.1 Flash for an open-weight release you can download now; pick Ling 3.1 Flash if a slightly higher Artificial Analysis score matters more than self-hosting.
- Gemini 3.8 Flash: Pick Gemini 3.8 Flash for image, audio and video input with published enterprise terms; pick Ling 3.1 Flash when text-only output at a lower list rate is enough.
- Claude Haiku 5.5: Pick Claude Haiku 5.5 for the lower list rate and higher index score on a generally available model; pick Ling 3.1 Flash only to trial its speed and reasoning style at no cost.