Gemini 4 Argon pricing, plans and limits

Google DeepMind's frontier reasoning model above the Gemini 3.8 Flash line, built for long coding, knowledge-work and cyber-defense jobs.

  • preview
  • proprietary
  • multimodal
  • Gemini 4 family
Google blog post announcing Gemini 4 Argon on September 30, 2026, by Koray Kavukcuoglu of Google DeepMind

As of 1 October 2026, only vetted security teams among the Fairwind Program's 650-plus partners can run Argon, so most buyers are planning for it rather than testing it. When access widens, it suits long codebase migrations and legal or finance research, where Vals AI scored it 68.9% on its cross-industry index.

Gemini 4 Argon is Google DeepMind's frontier reasoning model, scoring 53 with high reasoning on the Artificial Analysis composite of ten agentic, coding and knowledge tests. It reads text, images, audio and video, writes text, and stands out for calibration: its AA-Omniscience hallucination rate is 15%, so it usually declines rather than guesses.

Where it sits

  • $4.00/M$ per 1M tokensBlended price (3:1)Lower is better#48 / 75peer median $2.00/Mvendor price, checked by HokAI

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: Google DeepMind · Family: Gemini 4

More about Google DeepMind on HokAI

Context window: 1,000,000 tokens · Max output: 1,000,000

Input modalities: text, image, audio, video · Output: text

About Gemini 4 Argon

Gemini 4 Argon is the frontier reasoning model from Google DeepMind, announced on 30 September 2026 by Koray Kavukcuoglu, Google's Chief AI Architect. It is the first model in the Gemini 4 family, and Artificial Analysis describes it as Google's first proprietary model above the Flash class in more than seven months, after Gemini 3.1 Pro. Google has not disclosed the architecture, parameter count or training cutoff, and as of 1 October no model card or public API model ID exists. Artificial Analysis, Arena and Vals AI all list a 1M-token context window, but Google itself has not stated one. The rest of the lineup sits on the Google DeepMind models page.

Access is staged. Argon is going first to vetted security teams in Google's Fairwind Program, who get a version with the cyber guardrails removed so they can run it for vulnerability research. Google says paid Gemini API customers, who get their keys through Google AI Studio, and Google AI Ultra subscribers come next, with no date given. Until then the model can be studied but not bought.

On Google's own evaluations, Argon scores 77.9% on DeepSWE v1.1, a long-horizon software engineering test, and 51.3% on Zapier's AutomationBench. These are vendor-run numbers. The independent picture from Artificial Analysis is close to Google's story but less flattering at the top: its composite index puts Argon level with GPT-6 Astra and Anthropic's Fable 5.1, and a point ahead of GPT-6.1 Sol, while Claude Opus 5.5 sits at 58. That is a 23-point jump over Gemini 3.1 Pro Preview on the same index. On the AA variant of AutomationBench it scores 78%, and on Terminal Bench 4 it scores 57%, behind Claude Sonnet 5.5.

What Argon has already done inside Google is the strongest evidence so far, though it is Google's account. Argon agents replaced 32,000 lines of SIMD code in the Rust port of the libgav1 video decoder, producing safe Rust that the vendor reports runs 2.7 times faster than the earlier port with identical output. Another group of agents worked through fleet-wide profiling data and freed more than 300 TiB of memory across Google's data centers, Google says. On the security side, Google reports that Wiz, running it through its Scan for Good program, found a critical flaw in hospital software that earlier frontier models had missed.

The non-obvious cost is token volume. Artificial Analysis measured an average of about 62,000 output tokens per Intelligence Index task, more than twice what GPT-6 Astra used, so its low per-token rate does not translate one for one into a cheap job. Teams comparing it with Gemini 3.8 Flash or the earlier Gemini 3.8 Flash Cyber should price a real task, not a token. Google pairs the release with safety work it describes in some detail: activation monitoring to spot misuse, chain-of-thought monitoring that can stop an agent mid-run, refusal training for chemical, biological, radiological, nuclear and offensive cyber requests under its Frontier Safety Framework, and adversarial training against indirect prompt injection.

Who should care now: security teams eligible for Fairwind, and engineering, legal and finance groups that will want long single-pass work once the API opens. Who should not wait for it: anyone who needs a model in production this quarter, a published model card, or audio output. For those, a generally available model, such as the ones behind the Gemini app, or a rival frontier model is the practical choice today. Our guide to choosing an LLM walks through that decision by job, the launch comparison of Grok 4.7, Claude Opus 5, GPT-6 Astra and Gemini 3.8 Flash shows where the previous Gemini stood, and our analysis of Google's Gemini line covers the consumer side. To compare like with like, browse other preview-stage models, models that accept video or models with contexts above 400K tokens.

Screenshots

Google blog post announcing Gemini 4 Argon on September 30, 2026, by Koray Kavukcuoglu of Google DeepMind

Pricing

Introductory API price announced by Google: $2 per 1M input tokens and $10 per 1M output tokens, with cached input 95% cheaper than fresh input (about $0.10 per 1M). After the introductory period Google lists $4 input and $20 output per 1M tokens; the $0.20 cached rate in the table applies Google's stated 95% cache discount to that price. Google has not given an end date for the promotion; Artificial Analysis reports it runs for at least one month. Nobody can pay these rates yet: as of 1 October 2026 the model is not on the Gemini API pricing page, and access is limited to Fairwind Program partners. No batch, free or long-context tiers have been announced.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.060$0.010$0.070
Support reply$0.0040$0.0030$0.0070
One coding agent run$0.400$0.200$0.600

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • One-million-token output limit: A single response can run to 1M output tokens, up from 64K on earlier Gemini models, so a long migration or report can finish in one trajectory.
  • Long Decode Continuation: A new Gemini API mechanism pauses a very long response and resumes it through follow-up requests, so extended reasoning does not hit a request timeout.
  • Cyber defense mode: Trained to find, validate and patch software vulnerabilities, it scores 68% on CWE-bench v1 in Google's results, and Fairwind partners can run it inside the CodeMender security agent.
  • Long video understanding: Google reports 91.7% on LVBench, a test of locating details across long videos, alongside chart and multi-document analysis.
  • Four input types: Takes text, images, audio and video as input and answers in text; Artificial Analysis lists speech among the supported inputs.

Pros

  • Artificial Analysis measured a 15% hallucination rate on AA-Omniscience, against 51% for GPT-6 Astra, so it says "I don't know" far more often than it invents an answer.
  • Human raters on Arena gave its high-reasoning setting a preliminary text score of 1525 after 4,942 votes, a sign that the writing reads well, not only that it benchmarks well.
  • Vals AI's independent index of finance, legal, tax and coding work puts it at 68.9%, which matters more for business buyers than coding benchmarks alone.

Cons

  • No public access: there is no model ID, no Gemini API pricing entry and no model card, and Google has not dated the wider rollout.
  • Factual accuracy on AA-Omniscience is 50%, 13 points under GPT-6 Astra, so its caution comes partly from answering less, not knowing more.
  • Terminal Bench 4 lands at 57%, below Claude Sonnet 5.5 at 64%, so terminal-heavy coding agents still have stronger options.
  • Verbose by design: long reasoning traces mean a task can cost more than the per-token rate suggests.

Benchmarks

  • LMArena Elo: 1525 independent · 01 Oct 2026 — Rating from blind human votes on which answer is better.
  • LMArena rank: #1 independent · 01 Oct 2026 — Position on the blind human-preference leaderboard; #1 is best.
  • AA Intelligence Index: 53 cited: Artificial Analysis · 01 Oct 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

How much do you pay for Gemini 4 Argon?

Launch pricing is $2 in and $10 out per million tokens, cached input is 95% off, and the rate moves to $4 and $20 when the promotion ends. On Artificial Analysis testing that works out to $1.99 per Intelligence Index task now and $3.98 at the standard rate, against $3.26 for GPT-6 Astra. Nobody outside the Fairwind Program can buy it yet, so budget at the standard rate.

Is Gemini 4 Argon better than GPT-6 Astra?

On the overall Artificial Analysis score they tie, so behaviour and availability decide it. Argon is far more cautious about inventing facts and does better on business-workflow automation, while Astra answers more factual questions correctly and uses fewer than half the output tokens per task. Astra is on sale today and Argon is not, which settles it for anyone shipping this quarter.

Can you get access to Gemini 4 Argon yet?

Only through the Fairwind Program, which Google says works with over 650 partners and prioritises governments, critical infrastructure operators and core technology platforms. Applicants pass background checks, must use phishing-resistant multi-factor authentication, and may give access only to internal security, incident response or penetration testing staff. The weights are closed, so there is no download or self-hosting route either.

Why is Google releasing Gemini 4 Argon to cyber defenders first?

The same skills that let it find and patch vulnerabilities could help an attacker, so Google is giving defenders a head start while it tightens safeguards. Fairwind partners and Google's own teams get it without cyber guardrails, and their feedback shapes the refusal and monitoring systems before a wider release. Google also says it is taking part in the US government's voluntary pre-release access process.

Who should wait for Gemini 4 Argon, and who should look elsewhere?

Wait for it if your work is long and document-heavy: large code migrations, legal drafting or financial research, where a single pass can now run far longer than before. Look elsewhere if you need terminal-driven coding agents today (Claude Sonnet 5.5 scores higher on Terminal Bench 4), audio replies, or a published model card for a procurement review.

Top Alternatives

  • GPT-6 Astra: Pick GPT-6 Astra for factual accuracy and something you can deploy today; pick Gemini 4 Argon once released if invented answers are the bigger risk.
  • Claude Opus 5.5: Pick Claude Opus 5.5 when you want the strongest all-round independent results today; pick Gemini 4 Argon for very long single-response output.
  • Claude Sonnet 5.5: Pick Claude Sonnet 5.5 for terminal-heavy coding agents; pick Gemini 4 Argon for video and long-document analysis.
  • Gemini 3.8 Flash: Pick Gemini 3.8 Flash for generally available, lower-cost Gemini work now; pick Gemini 4 Argon for the hardest long-horizon tasks when it opens.
  • Gemini 3.8 Flash Cyber: Pick Gemini 3.8 Flash Cyber if your team already has it through Google; pick Gemini 4 Argon for deeper vulnerability discovery once approved.

More AI Models on HokAI

Visit Gemini 4 Argon Official Page