Cohere Command A+ review, pricing and verdict

Cohere's flagship open-weights model, unifying Command A's reasoning, vision, and translation lines into one Apache 2.0 checkpoint for enterprise agents.

  • ga
  • open source
  • multimodal
  • Command A family

Command A+ suits enterprises wanting self-hosted or sovereign AI: it runs on modest GPU hardware at Cohere's recommended quantization and ships under a fully open license with no per-token lock-in. It replaces four separate Cohere models but trails today's frontier closed models on raw reasoning benchmarks.

Cohere Command A+ is an open-weights sparse mixture-of-experts language model that unifies Cohere's separate reasoning, vision, and translation models into one checkpoint. Artificial Analysis measured it at 76% on GPQA Diamond and scored it 37 on its Intelligence Index; the model is released under a fully permissive Apache license for self-hosted or Cohere-managed deployment.

Where it sits

  • 281 tok/stokens/sOutput speedHigher is better#8 / 49peer median 90 tok/scited: Artificial Analysis
  • 76%% correctGPQA DiamondHigher is better#35 / 50peer median 88.3%per source, see benchmark scores

In the bottom third on GPQA Diamond (rank 35 of 50), and one of 76 whose vendor states it does not train on customer data.

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: Cohere · Family: Command A

More about Cohere on HokAI

Context window: 128,000 tokens · Max output: 64,000

Input modalities: text, image, tool-calls · Output: text, tool-calls

About Cohere Command A+

Cohere Command A+ is a sparse mixture-of-experts foundation model released by Cohere on May 20, 2026 under the model ID command-a-plus-05-2026. It has 218 billion total parameters with 25 billion active per token, architected to run on far less hardware than its total parameter count suggests. Command A+ unifies capabilities that were previously split across four separate Cohere models, Command A, Command A Reasoning, Command A Vision, and Command A Translate, into a single scalable model, making it Cohere's first model to combine multimodal reasoning and machine translation in one checkpoint. It sits at the top of Cohere's Command lineup, above Command R+ and the smaller Command R7B.

On benchmarks, Artificial Analysis gives Command A+ an Intelligence Index of 37 and measured 76% on GPQA Diamond, while Cohere reports 63% on MMMU-Pro and a Terminal-Bench Hard gain from 3% to 25% over Command A Reasoning. Every figure below was re-checked against its source on 30 September 2026:

| Benchmark | Command A+ | Earlier Cohere model | Source | |---|---|---|---| | AA Intelligence Index | 37 | not given | Artificial Analysis | | GPQA Diamond | 76% | not given | Artificial Analysis | | Humanity's Last Exam | about 11% | not given | Artificial Analysis | | SciCode | about 38% | not given | Artificial Analysis | | AA-Omniscience non-hallucination | 86% | not given | Artificial Analysis | | Terminal-Bench Hard | 25% | 3% (Command A Reasoning) | Cohere | | tau-squared-Bench Telecom | 85% | 37% (Command A Reasoning) | Cohere | | MMMU | 75.1% | 65.3% (Command A Vision) | Cohere | | MMMU-Pro | 63% | not given | Cohere | | MathVista | 80.6% | 73.5% | Cohere | | CharXiv reasoning | 52.7% | 46.9% | Cohere | | Terminal-Bench 2.1 | 17.6% | not given | Vals AI |

Read the two sources differently. Cohere's figures compare the model with its own predecessors and come from Cohere's launch post; Artificial Analysis and Vals AI ran their own harnesses. Vals AI lists an average accuracy of 32.69% across its suite, including 61.42% on LegalBench and 9.04% on Finance Agent v2. Inside Cohere's North workspace, Cohere reports agentic question answering up 20% and spreadsheet analysis up 32% over Command A Reasoning, which are internal evals with no public harness.

Command A+ supports a 128,000-token input context window with a 64,000-token maximum generation length. Cohere reports output throughput up to 63% higher and time-to-first-token 17% lower than Command A Reasoning at the same quantization, with the W4A4 4-bit build adding a further 47% speed gain and 13% latency cut. Artificial Analysis measured about 281 output tokens per second on Cohere's API in pre-release testing.

The model accepts text, image, and tool-use input and produces text and tool-use output. Language coverage grew from 23 to 48 languages, and Cohere's new tokenizer uses 20% fewer tokens for Arabic, 18% fewer for Japanese and 16% fewer for Korean. Function calling and agentic tool use are core design targets, which is where the largest benchmark gains in the table sit.

Command A+ is released under a full Apache 2.0 license, and Cohere publishes the weights in three formats through Hugging Face, under CohereLabs/command-a-plus-05-2026: BF16 (4x B200 or 8x H100 GPUs), FP8 (2x B200 or 4x H100), and W4A4 (1x B200 or 2x H100, Cohere's recommended default). Cohere's model card states benchmark quality differences across the three are negligible. Cohere has not published a per-token price for Command A+; it is positioned for self-hosting or Cohere's Model Vault managed deployment, which keeps data inside a customer's own environment and is marketed at regulated industries needing sovereign or air-gapped AI.

Command A+ carries Cohere's two-mode safety configuration: contextual mode for wide-ranging interactions that still rejects clearly harmful or illegal requests, and strict mode that avoids violence, sexual content, and profanity entirely. Cohere has not published a third-party red-team partner list or safety benchmark numbers for this model.

Who should look elsewhere: teams whose work depends on hard science or agentic coding scores, where the table shows the gap, or on documents past 128K tokens. Mistral Large 3 offers a 256K-token window and a published API rate, DeepSeek V4 leans further into agentic coding, Hy4 Preview is a larger open MoE checkpoint, and Gemini 3.1 Flash-Lite is a hosted model Artificial Analysis set beside it in its launch comparison. Teams without GPUs will also find a metered API simpler than a Model Vault contract. For the wider field, see our local coding model picks, how to choose an LLM, and the full model directory.

Pricing

Command A+ has no confirmed public per-token API price as of mid-2026; it is released open-weights under Apache licensing, priced for self-hosting or a Cohere Model Vault contract instead of metered API billing. Infrastructure is the real cost: 24-hour on-demand cloud rental runs about $55 for a single B200 GPU or $72 for two H100 GPUs at the recommended W4A4 quantization. For a hosted reference point, Cohere's older Command R+ API charges $2.50 per 1M input tokens and $10.00 per 1M output tokens.

TierRate
Self-hosted (Apache 2.0)Free weights, infrastructure cost only (min 1x B200 or 2x H100 at W4A4)

Key Features

  • Unified Reasoning, Vision, and Translation: Merges four previously distinct Cohere models, spanning reasoning, vision, and translation workloads, into one checkpoint that handles multimodal reasoning and machine translation together.
  • Sparse Mixture-of-Experts Routing: Routes each token through a small subset of specialized experts, letting the model operate at frontier scale while running on a fraction of the compute a similarly sized dense model would need.
  • Expanded Multilingual Coverage: Covers all official EU languages, with Cohere reporting measurable tokenization efficiency gains for Arabic, Japanese, and Korean text versus the prior Command A Reasoning model.
  • 4-bit W4A4 Quantized Deployment: Cohere's recommended default deployment mode, which Cohere reports keeps benchmark quality loss negligible while trimming GPU requirements down to a single modest card.
  • Fully Open License: Weight files are downloadable from Hugging Face with no usage restrictions, letting teams audit, fine-tune, or run the model fully air-gapped.

Pros

  • Ships under a fully open license with weights on Hugging Face, so there is no per-token billing lock-in or dependency on Cohere's API uptime.
  • Posts large agentic gains over its predecessor: Terminal-Bench Hard rose from 3% to 25% and tau-squared-Bench Telecom from 37% to 85%.
  • Runs at Cohere's recommended W4A4 quantization on far more modest hardware than its total parameter count implies, unusually light for a model this size.

Cons

  • Its composite score from Artificial Analysis still trails frontier closed models such as GPT-5-series and Claude Opus on composite reasoning.
  • There is no confirmed low-cost hosted per-token API; production access effectively requires self-hosting or a Cohere Model Vault contract.
  • The 128,000-token context window trails several 2026 rivals now offering windows past 1 million tokens.

Benchmarks

  • MMMU-Pro: 63% vendor-reported · 30 Sep 2026 — College-level questions that need reading images and diagrams, % correct.
  • GPQA Diamond: 76% cited: Artificial Analysis · 30 Sep 2026 — PhD-level science questions that are hard to search for, % correct.
  • Terminal-Bench 2.1: 17.6% independent · 30 Sep 2026 — Multi-step tasks completed in a real command line, % solved.
  • Humanity's Last Exam: 11% cited: Artificial Analysis · 30 Sep 2026 — Expert-written questions across many fields, % correct.
  • AA Intelligence Index: 37 cited: Artificial Analysis · 30 Sep 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
  • Output speed: 281 tok/s cited: Artificial Analysis · 30 Sep 2026 — Median tokens written per second as measured by Artificial Analysis.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

What are Cohere Command A+'s pricing plans in 2026?

There is no published per-token price for Command A+ itself since Cohere ships it as open weights rather than a metered API. Expect to pay for compute instead: renting a single high-end GPU for a day runs roughly $55, and a two-GPU setup runs about $72, both at Cohere's recommended 4-bit quantization. Cohere's legacy Command R+ API, priced at $2.50 input and $10.00 output per 1M tokens, is the closest hosted comparison if a per-token number is what you need.

Is Cohere Command A+ free to use?

Yes: Command A+ is released under a permissive open-weights license with no usage restrictions, so the model itself costs nothing to download or run commercially. The only real cost is GPU infrastructure, since Cohere has not published a free hosted API tier for it the way it has for some smaller Command models.

What are the best alternatives to Cohere Command A+?

Mistral Large 3 and DeepSeek V4 are the closest open-weights alternatives if self-hosting matters: Mistral Large 3 trades some scale for a smaller footprint, while DeepSeek V4 leans further into agentic coding benchmarks. Hy4 Preview is worth a look for teams that want an even larger open MoE checkpoint with an Apache-licensed release.

Is Cohere Command A+ better than Mistral Large 3?

Command A+ and Mistral Large 3 are both fully open, self-hostable MoE models with no per-token API lock-in, but they target different jobs. Mistral Large 3 offers a wider 256K-token context window at a lower published API rate, while Command A+ adds native multimodal reasoning, machine translation, and support for far more languages in one checkpoint. Teams choosing between them should weigh long-document handling against unified multimodal and multilingual coverage.

How long does it take to get going with Cohere Command A+?

Getting started means downloading the weights from Hugging Face under CohereLabs/command-a-plus-05-2026 and picking a quantization: most teams start with the W4A4 4-bit build, which Cohere says keeps benchmark quality loss negligible on comparatively modest self-hosted hardware. Teams that would rather skip infrastructure entirely can request access to Cohere's managed Model Vault deployment instead.

Top Alternatives

  • Mistral Large 3: Pick Mistral Large 3 for a wider context window and lower published API pricing; pick Command A+ for unified multimodal reasoning, translation, and broader language coverage in one checkpoint.
  • DeepSeek V4: Pick DeepSeek V4 if agentic coding benchmark score is the deciding factor; pick Command A+ if you need native vision and translation alongside reasoning in a single model.
  • Hy4 Preview: Pick Hy4 Preview for a larger open-weight checkpoint from Tencent; pick Command A+ for Cohere's enterprise-focused Model Vault deployment path and multilingual tooling.

More AI Models on HokAI

Visit Cohere Command A+ Official Page