Ling 3.0 Flash Fin review, pricing and limits

InclusionAI's finance-specialized fine-tune of Ling 3.0 Flash, tuned for source-grounded investment research, valuation and spreadsheet workflows.

  • ga
  • open source
  • chat
  • Ling 3.0 family
checked

Ling 3.0 Flash Fin is InclusionAI's finance research model built on Ling 3.0 Flash, running on a 124B/5.1B-parameter MoE architecture with a 262,144-token context window. It costs $0 per million tokens on Novita AI and OpenRouter's free tier, and suits teams that want an open, self-hostable finance model over an unverified benchmark claim.

Ling 3.0 Flash Fin is InclusionAI's finance-domain fine-tune of Ling 3.0 Flash, released 27 August 2026 with 124 billion total and about 5.1 billion active MoE parameters. It serves a 262,144-token context window and native tool calling, but InclusionAI has not published numeric benchmark scores for it, only a chart image.

Where it sits

  • $0/M$ per 1M tokensBlended price (3:1)Lower is better#1 / 59peer median $2.00/Mvendor price, checked by HokAI
  • --tokens/sOutput speedHigher is better-- / 33peer median 90 tok/scited: Artificial Analysis
  • --% solvedSWE-bench VerifiedHigher is better-- / 26peer median 78.3%per source, see benchmark scores
  • --% correctGPQA DiamondHigher is better-- / 41peer median 86.9%per source, see benchmark scores

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: InclusionAI · Family: Ling 3.0

More about InclusionAI on HokAI

Context window: 262,144 tokens · Max output: 32,768

Input modalities: text, tool-calls · Output: text, tool-calls

About Ling 3.0 Flash Fin

Ling 3.0 Flash Fin is InclusionAI's first finance-domain fine-tune, released on 27 August 2026 as a continued-training extension of the general-purpose Ling-3.0-flash model. InclusionAI, the AI research lab inside Ant Group, built it together with unnamed financial institutions and domain experts, applying additional training on financial data on top of the same Ling 3.0 Flash backbone rather than starting a new architecture from scratch. It sits in the Ling family alongside a health-domain sibling, Ling-3.0-flash-Sante, showing InclusionAI's strategy of shipping several narrow domain fine-tunes off one shared base model instead of one general-purpose flagship. Architecturally it inherits Ling 3.0 Flash's BailingMoeV3ForCausalLM design: a sparse mixture-of-experts network with 512 routed experts and 8 active per token plus one shared expert, a native multi-token-prediction head, and a hybrid linear-attention stack that alternates Kimi Delta Attention and multi-head latent attention in a 5:1 ratio. Total parameter count is 124 billion, with roughly 5.1 billion active per token, an activation ratio InclusionAI compressed to 1/64 versus 1/32 in the prior Ling generation specifically to cut inference cost while keeping quality. InclusionAI has not published numeric benchmark scores for this model. Its own model card lists seven finance-specific evaluation suites it says it ran: FinFIRST for source-grounded retrieval, FinSearchComp Verified, FinCRAFT, Finance Agent, APEX-Agents for long-horizon execution, SpreadsheetBench for valuation and spreadsheet operations, and tau3-Banking for banking workflows. The card states the model is 'competitive with both similarly sized models and substantially larger general-purpose models,' but that claim is illustrated only with an unlabeled chart image, not a table of numbers, and no independent evaluator (Artificial Analysis, LMArena, or otherwise) has published a score for this specific fine-tune either. The model serves a 262,144-token context window with up to 32,768 tokens of output, handles text input and output only (no vision, audio, or video), and supports OpenAI-compatible tools and tool_choice parameters for function calling. Extended 'thinking' mode is enabled by default rather than opt-in, and InclusionAI's own model card recommends running inference at temperature 1.0, top_p 0.95 and top_k 20 for best results. As of September 2026 the model is hosted for free, at $0 per million input and output tokens, on Novita AI and via OpenRouter's free-tier listing, and Novita's own page explicitly notes that rate is not a permanent commitment. Other aggregators list a standard, non-free tier around $0.06 per million input tokens and $0.18 per million output tokens on some hosts, though pricing for this specific fine-tune (as opposed to the base Ling 3.0 Flash) is not confirmed on every provider. It is also listed on Fireworks AI, Baseten, DeepInfra, and Vercel's AI Gateway, and the raw weights are on Hugging Face as a BF16 safetensors checkpoint. Being MIT-licensed and open-weight, it can be self-hosted. A community GGUF conversion (bartowski/Ling-3.0-flash-Fin-GGUF) is available for llama.cpp, with a Q4_K_M quantization around 78GB recommended as the size-versus-quality balance; smaller quantizations start near 32GB. On a single GPU with less VRAM than that, llama.cpp's MoE-offload flags let routed experts run from system RAM instead, at a real speed cost. InclusionAI frames the target use case as financial research workflows: information retrieval across financial documents, evidence review, valuation calculation, spreadsheet modeling, and report preparation, with an emphasis on source-grounded, multi-document answers rather than open-ended chat. The model card is explicit about where that stops: 'as our first finance-enhanced release, Ling-3.0-flash-Fin still requires further validation in complex, long-horizon workflows,' and its outputs 'do not constitute investment advice' and need professional review before use in any real financial decision. InclusionAI has not published a training data cutoff date or a dedicated safety system card for this model. The base Ling 3.0 Flash pretraining checkpoint's own documentation notes it is not intended for direct end-user chat or safety-critical use without further alignment, implying the production Flash and Fin checkpoints do carry some post-training alignment, though InclusionAI does not name a specific method (no disclosed Constitutional AI, RLAIF or equivalent framework) or any red-teaming partners, unlike the system-card practices of Western frontier labs.

Pricing

Both Novita AI and OpenRouter's free tier list this model for nothing as of September 2026, but Novita explicitly flags that rate as promotional, not a permanent commitment. A handful of other aggregator listings show a paid tier instead, priced well under a dime per million tokens either direction, though that rate is not independently confirmed for this exact fine-tune on every provider. As an MIT-licensed open-weight model it can also be self-hosted at pure infrastructure cost, with no InclusionAI fee.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0$0$0
Support reply$0$0$0
One coding agent run$0$0$0

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • Finance-domain continued training: Extends Ling 3.0 Flash through continued training on financial data developed with financial institutions and domain experts, rather than a generic instruction fine-tune.
  • 262K-token context window: Handles a long context with a five-figure output ceiling, enough to hold multiple financial filings or reports in a single request.
  • Thinking mode by default: Ships with extended reasoning turned on by default, and InclusionAI's model card lists specific recommended decoding settings for best results.
  • Native tool calling: Supports OpenAI-compatible tools and tool_choice parameters on hosted endpoints, letting it call external functions for retrieval, calculation or spreadsheet operations mid-task.
  • MIT-licensed open weights: Released under the MIT license as a BF16 safetensors checkpoint with community GGUF quantizations, so it can be self-hosted without per-token API costs.

Pros

  • Free to use on multiple hosted endpoints (Novita AI, OpenRouter's free tier) as of September 2026, with no per-token cost on those hosts.
  • 262,144-token context window comfortably fits long financial documents and multi-document research tasks in one request.
  • MIT license and public GGUF quantizations mean it can be self-hosted with no ongoing API bill, unlike closed finance-tuned models from Western vendors.

Cons

  • No public, numeric benchmark scores for any of its seven finance evaluation suites, only an unlabeled chart image on the model card, so its real-world financial accuracy is unverified by any independent source.
  • Text-only: cannot read images, scanned PDFs, or charts directly, a real gap for financial workflows that often start from scanned filings.
  • InclusionAI itself calls it a first-generation release still needing validation in complex, long-horizon workflows, and its free hosted pricing is explicitly not a permanent commitment.

Benchmarks

    A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

    Frequently Asked Questions

    How much does Ling 3.0 Flash Fin cost per 1M tokens?

    It is free on Novita AI and OpenRouter's free-tier listing as of September 2026, at $0 per million input and output tokens, though Novita's own page states that rate is not a permanent commitment. Some other hosts list a standard, non-free tier around $0.06 per million input tokens and $0.18 per million output tokens. Being MIT-licensed and open-weight, it can also be self-hosted at pure infrastructure cost with no fee to InclusionAI.

    Does Ling 3.0 Flash Fin publish benchmark scores?

    InclusionAI's own model card lists seven finance-specific evaluation suites it says it ran, including FinFIRST, FinSearchComp Verified, FinCRAFT, Finance Agent, APEX-Agents, SpreadsheetBench and tau3-Banking, but only shows an unlabeled chart image rather than a table of numeric scores. No independent evaluator such as Artificial Analysis or LMArena has published a score for this specific fine-tune either, so its accuracy claims are currently unverifiable.

    Is Ling 3.0 Flash Fin open source?

    Yes. It is released under the MIT license as a BF16 safetensors checkpoint on Hugging Face, with a community GGUF conversion available for self-hosting via llama.cpp. It is also available through hosted API endpoints on OpenRouter, Novita AI, Fireworks AI, DeepInfra, Baseten and Vercel's AI Gateway.

    Does Ling 3.0 Flash Fin train on user data?

    InclusionAI has not published a data retention or training-on-inputs policy for this model directly. Since it has no single vendor API, retention depends on which hosting provider is used (OpenRouter, Novita AI, Fireworks AI, DeepInfra, Baseten or Vercel AI Gateway), each of which sets its own policy, or on infrastructure the user controls if self-hosted.

    Who is Ling 3.0 Flash Fin best for and who should avoid it?

    It suits fintech and research teams that want a free or self-hostable open-weight model for source-grounded financial document research with a long, 262,144-token context window. Teams needing regulated investment advice, an independently benchmarked model, or direct image and scanned-document understanding should look elsewhere, since InclusionAI's own card disclaims investment advice and the model is text-only.

    More AI Models on HokAI

    Visit Ling 3.0 Flash Fin Official Page