K-EXAONE 2.0 review, pricing and limits

LG AI Research's open-weight flagship, a frontier-scale mixture-of-experts model built for Korean sovereign AI and self-hosted agent work.

  • ga
  • open source
  • chat
  • EXAONE family

Choose it if you need a permissively licensed model you can host yourself, with a 262,144-token window and strong long-context retrieval in its own comparison table. Skip it when coding or math strength is the priority, because DeepSeek V4 Pro and GLM score higher on most of those rows.

What changed

  • · New feature FP8, NVFP4 and DSpark checkpoints added to the collection (K-EXAONE-2.0-750B-A37B) source
  • · Release K-EXAONE 2.0 published as a 750B open-weight model (K-EXAONE-2.0-750B-A37B) source

All changes across HokAI this week

K-EXAONE 2.0 is LG AI Research's open-weight mixture-of-experts language model, released on July 31, 2026 under the Apache License 2.0. It has 750 billion total parameters, activates 37 billion per token, and reports 68.2% on SWE-bench Verified in its own model card.

Where it sits

  • 68.2%% solvedSWE-bench VerifiedHigher is better#25 / 33peer median 77.2%per source, see benchmark scores
  • 82.2%% correctGPQA DiamondHigher is better#34 / 51peer median 87.8%per source, see benchmark scores

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: LG AI Research · Family: EXAONE

More about LG AI Research on HokAI

Context window: 262,144 tokens

Input modalities: text, tool-calls · Output: text, tool-calls

About K-EXAONE 2.0

K-EXAONE 2.0 is the newest text model from LG AI Research, published on Hugging Face on July 31, 2026. It is a mixture-of-experts language model with 750 billion total parameters, of which 37 billion are active per token, and it is the second model the lab built for South Korea's Sovereign AI Foundation Model Project (see all LG AI Research models in the HokAI model directory). The first K-EXAONE had 236 billion parameters (23 billion active). LG AI Research made the jump by upcycling that model, widening and deepening it, then running continual pretraining, difficulty-focused mid-training and post-training, and it describes the result as the largest foundation model built in Korea so far.

The architecture has 78 main layers (2 dense layers followed by 76 sparse ones) plus one multi-token-prediction layer. Each sparse layer holds 256 routed experts and one shared expert, with 8 routed experts active. Attention mixes sliding-window layers (128 and 4,096 token windows) with periodic global layers, using 64 query heads and 8 key-value heads. The vocabulary is 153,600 tokens, the context length is 262,144 tokens (a window in the 128K to 400K bracket) and the stated knowledge cutoff is the second quarter of 2025.

On the benchmark table in its model card, K-EXAONE 2.0 scores 83.5 on MMLU-Pro, 82.2 on GPQA-Diamond, 92.3 on AIME 2026, 68.2 on SWE-bench Verified and 43.8 on Terminal-Bench 2.1, with 18.3 on Humanity's Last Exam. Those figures are vendor-reported; no independent evaluation was found when this page was written. The same table puts Qwen3.5, GLM-5.1 and DeepSeek V4 Pro ahead on most knowledge and coding rows, so this is a competitive open model rather than the leader on those axes. The lab says its average over 24 benchmarks rose from 63.3 for the first K-EXAONE to 70.1, and that three coding and agentic coding tests improved by about 30%.

Long-context retrieval is where the card makes its strongest case: 94.4 on OpenAI-MRCR and 89.6 on Ko-LongBench, and 56.2 on AA-LCR, where all three rivals in the table score higher. Thinking mode is on by default, and a preserve_thinking option carries earlier reasoning across turns, which the card recommends for agent runs and deep research. Tool calling is supported in both serving engines. Agentic tool use is modest: 14.2 on the tau3 Banking test.

Language coverage grew from six to ten: Korean, English, Spanish, German, Japanese, Vietnamese, French, Italian, Polish and Portuguese. Korean is the lab's home market, but the model's own Korean knowledge scores are not the highest in the table: it posts 69.1 on KMMLU-Pro against 75.8 for GLM-5.1. Buyers choosing it for Korean quality alone should test their own prompts first, and can compare HyperCLOVA X from Naver (more in Naver's model list) and Solar Pro 4 from Upstage (Upstage's models). It handles text only.

The weights are published as BF16, FP8, NVFP4 and a DSpark speculative-decoding build, all under the Apache License 2.0, which places it among the open-source licensed models and the generally available releases. The card's serving examples use SGLang and vLLM from LG-linked forks on two nodes of eight NVIDIA H200 GPUs, and says MTP and DSpark speculative decoding speed up generation by roughly three to five times. FriendliAI lists a dedicated-endpoint option, and a browser demo runs at k.exaone.ai. No per-token API price was published on any page checked, so cost depends on your own GPU budget.

For safety, the lab reports 99.8 on KGC-Safety and 89.5 on ROK-Fortress, two tests that cover Korean and global safety standards and geopolitically sensitive contexts, with an average of 94.6 across the two. As with every language model, the card warns of biased or inappropriate output and stale knowledge. It is a strong fit for organizations that need a permissively licensed, self-hostable model with Korean-market and sovereignty requirements, and a weaker fit for teams that want the top coding model, where DeepSeek V4 Pro and Z.ai's GLM line (HokAI lists GLM-5.3) post higher scores in the card.

Pricing

LG AI Research publishes no per-token API price for K-EXAONE 2.0 on the pages HokAI checked. The cost is the GPU fleet the card describes for self-hosting. Per-token alternatives to price against include [DeepSeek V4 Pro](/hub/models/deepseek-v4-pro) and [Kimi K3](/hub/models/kimi-k3).

Key Features

  • Mixture-of-experts design: 256 routed experts plus one shared expert per sparse layer, with 8 routed experts firing for each token.
  • Hybrid attention: Blocks of three sliding-window layers and one global layer, plus a NoPE global layer, to keep long prompts affordable.
  • Switchable thinking mode: Reasoning is on by default and can be turned off per request, with a preserve_thinking flag for multi-turn agent runs.
  • Built-in speculative decoding: MTP and DSpark builds add speculative decoding that the model card says accelerates output several-fold.
  • Apache-licensed weights in four builds: Full-precision, FP8, NVFP4 and DSpark checkpoints are public on Hugging Face for commercial use.

Pros

  • Permissive Apache license on frontier-scale weights, so a company can host, fine-tune and ship it without a usage agreement.
  • Leads the comparison table on both of the lab's safety tests, which matters for public-sector and regulated buyers in Korea.
  • Quantized FP8 and NVFP4 builds are published by the vendor, which lowers the memory bill for self-hosting.

Cons

  • Trails the Chinese open models in its own table on coding, math and knowledge, including 69.1 on the Korean KMMLU-Pro test.
  • Needs multi-node GPU hardware, since the card's serving example spans two eight-GPU nodes.
  • Quickstart relies on forked inference engines and no per-token hosted price is published, which raises setup effort.

Benchmarks

  • MMLU-Pro: 83.5% vendor-reported · 12 Oct 2026 — A harder version of the 57-subject knowledge exam, % correct.
  • AIME 2026: 92.3% vendor-reported · 12 Oct 2026 — Competition-level maths problems from the 2026 exam, % solved.
  • GPQA Diamond: 82.2% vendor-reported · 12 Oct 2026 — PhD-level science questions that are hard to search for, % correct.
  • SWE-bench Verified: 68.2% vendor-reported · 12 Oct 2026 — Real GitHub issues fixed end to end, % solved.
  • Terminal-Bench 2.1: 43.8% vendor-reported · 12 Oct 2026 — Multi-step tasks completed in a real command line, % solved.
  • Humanity's Last Exam: 18.3% vendor-reported · 12 Oct 2026 — Expert-written questions across many fields, % correct.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

What does K-EXAONE 2.0 actually cost?

The weights are free to download, so the real cost is hardware: the model card's own serving example uses two nodes of eight NVIDIA H200 GPUs. The vendor lists no per-token API rate, and the FriendliAI page checked on 2026-10-12 offered a dedicated-endpoint deploy option with no token rate shown. Compare that capital cost with per-token models such as [DeepSeek V4 Pro](/hub/models/deepseek-v4-pro) before committing.

How does K-EXAONE 2.0 compare with DeepSeek V4 Pro on benchmarks?

In the model card's table, DeepSeek V4 Pro (max) is ahead on SWE-bench Verified (80.6) and GPQA-Diamond (90.1). K-EXAONE leads on the lab's safety tests, 99.8 on KGC-Safety against 82.8, and edges ahead on OpenAI-MRCR retrieval (92.9 for DeepSeek). All figures are vendor-reported, so rerun a sample of your own tasks.

Is K-EXAONE 2.0 open source?

The weights are public under the Apache License 2.0, which permits commercial use, and the vendor also ships FP8 and NVFP4 builds. The model card does not publish the training data, so strictly speaking it is an open-weight release with a permissive license rather than a fully open pipeline.

Does K-EXAONE 2.0 train on my data?

When you self-host the weights, prompts stay on your own infrastructure and LG AI Research is not in the loop. HokAI did not review the data terms of the k.exaone.ai demo or the FriendliAI endpoint, so read those policies before sending sensitive text. Technical support is listed at contact_us@lgresearch.ai.

Who should use K-EXAONE 2.0, and who should skip it?

It suits organizations that want a domestic, permissively licensed model they can run inside their own network, especially in Korea. Skip it for coding agents where raw SWE-bench strength decides the choice, and for small teams without a GPU cluster, who may prefer the smaller self-hostable variants of [HyperCLOVA X](/hub/models/hyperclova-x).

Top Alternatives

  • DeepSeek V4 Pro: Pick DeepSeek V4 Pro if coding and math scores decide it; pick K-EXAONE 2.0 for a smaller active footprint and the Apache license.
  • HyperCLOVA X: Pick HyperCLOVA X if you want smaller self-hostable Korean models; pick K-EXAONE 2.0 when you have the GPUs for frontier scale.

More AI Models on HokAI

Visit K-EXAONE 2.0 Official Page