Grok 4.6 reportedly keeps the same 1.5 trillion-parameter base xAI used for Grok 4.5, so this release's gains sit entirely in post-training, not scale. Treat it as a cheaper, refined version of 4.5 rather than a new tier, since xAI's roadmap already points to a bigger Grok 4.7 within weeks.
Grok 4.6 is xAI's 2026 update to its flagship reasoning model, scoring 61 on the Artificial Analysis Intelligence Index, tied for the top spot among 184 tracked models, and 69.9% on CursorBench v3.2. It keeps Grok 4.5's 500,000-token context window, concentrating this release's gains in post-training rather than model scale.
Provider: xAI · Family: Grok 4
Context window: 500,000 tokens
Input modalities: text, image, tool-calls · Output: text, tool-calls
About Grok 4.6
Grok 4.6 is xAI's flagship reasoning model, released in August 2026. xAI merged into SpaceX in an all-stock deal announced February 2026 and now trades as SPCX after a June 2026 Nasdaq listing. Rather than scaling up, xAI held its base model constant: multiple independent trackers report Grok 4.6 reuses the same 1.5 trillion-parameter V9 foundation as Grok 4.5, with xAI's release notes citing a longer post-training run using curated model-generated reasoning data, engineering data, and agentic RL tasks instead of a larger pretrain. xAI has not officially disclosed a parameter count or architecture type for any Grok 4.x model. On the Artificial Analysis Intelligence Index, an independent cross-vendor benchmark, Grok 4.6 scores 61, tying GPT-5.6 Sol Max for the top spot among 184 tracked models. xAI's launch benchmarks lean agentic: 69.9% on CursorBench v3.2, 65.9% on DeepSWE v1.1, 61.3% on FrontierCode v1.1 Extended, 57.5% on APEX-Agents, 56.4% on APEX-SWE, 26% on Terminal-Bench v3.0, a GDPVal-AA v2 score of 1753, an AA-Briefcase score of 1577, and 15.8% on Harvey LAB (Vals), a legal-reasoning eval. As with Grok 4.5, xAI did not publish GPQA Diamond, AIME 2025, MMLU-Pro, or SWE-bench Verified scores at launch. The context window holds at 500,000 tokens, unchanged from Grok 4.5 and half of Grok 4.3's 1 million. xAI has not published a maximum output-token figure for Grok 4.6. Artificial Analysis independently measured 67.6 output tokens per second on the high-reasoning-effort configuration. Input modalities cover text and images up to 20MiB per file in jpg, jpeg, or png format; output is text only, so image or video generation needs a separate model. Grok 4.6 exposes an adjustable reasoning_effort parameter with four depths (low, medium, high, xhigh), native function calling, structured output, and an integrated web-search tool, so agent loops do not need a separate retrieval layer. Per-token pricing stayed flat versus Grok 4.5 for prompts that stay under a set length, then steps up to a long-context rate beyond it, the same two-tier mechanic Grok 4.5 used (see price_notes and the pricing FAQ for exact figures). xAI also confirmed a faster variant priced at roughly twice the standard rate, though it had not published exact figures or a separate model slug as of launch. At launch, Grok 4.6 is reachable through xAI's direct API (api.x.ai, OpenAI-compatible, Bearer-token auth, Python via xai_sdk, TypeScript via the Vercel AI SDK) and through OpenRouter. It shipped live the same day inside Cursor and xAI's own Grok Build, with both getting double included usage for the first week. It is not yet confirmed on AWS Bedrock, Azure AI Foundry, or Google Vertex, which still carry the older Grok 4.3. Safety training combines supervised fine-tuning with reinforcement learning from human feedback, verifiable rewards, and model-based grading, tuned to refuse requests showing clear intent toward severe harm or criminal activity. A layered stack of runtime input and topical filters adds controls for CSAM, self-harm, and biological or chemical weapons pathways, plus cyber-specific input safety controls, while the system prompt is tuned to avoid over-refusal on benign or hypothetical discussions. Enterprise customers can enable Zero Data Retention; xAI states SOC 2 Type 2 compliance and offers a BAA for healthcare customers on the ZDR-enabled API. Grok 4.6 fits teams already building agent loops in Cursor or Grok Build, or anyone who wants a top-tier Artificial Analysis Intelligence Index score without paying more than Grok 4.5 cost. It is a weaker pick for teams needing GPQA, AIME, or MMLU-Pro comparability for procurement, native audio or video understanding, or a managed-cloud endpoint on day one. xAI's roadmap points to Grok 4.7, a ~2.1 trillion-parameter step up expected weeks later; Grok 4.6 is best understood as a post-training refinement of Grok 4.5, not a new capability tier.
Pricing
Below a 200,000-token prompt, Grok 4.6 bills $2/$6 per million input/output tokens, with a quarter-price $0.50 per million rate on cached input (a 75% discount). Cross that threshold and the entire request shifts to the long-context band at $4/$12 per million, with cached input at $1, not just the tokens over the line. xAI has also flagged a faster variant at roughly double the cost, without yet publishing its exact rate or a distinct model slug.
Key Features
- Tiered Pricing by Prompt Size: A standard per-token rate applies below a set prompt-length threshold; cross it and pricing roughly doubles for that entire request, not just the excess.
- Four Reasoning Depths: The reasoning_effort parameter accepts low, medium, high, or xhigh, trading latency for accuracy on a per-request basis.
- Built-in Web Search and Tool Calling: Ships with an integrated web-search tool and native function calling, so agent loops can pull live data without a separate retrieval layer.
- 75% Cached-Input Discount: Cached input tokens cost a quarter of the standard input rate, rewarding agent loops that reuse the same system prompt across turns.
- Live in Cursor and Grok Build at Launch: Ships same-day inside Cursor and xAI's coding tool Grok Build; both get double the included usage during launch week.
Pros
- Ranks at the very top of the independent, cross-vendor Artificial Analysis Intelligence Index for frontier models.
- Keeps Grok 4.5's per-token pricing despite the capability upgrade, and rewards reused system prompts with a cached-input discount.
- Adjustable reasoning_effort plus a built-in web-search tool, so agent loops skip a separate retrieval layer.
Cons
- Hasn't published scores on the standard academic reasoning benchmarks most rival frontier models report, making head-to-head comparison harder.
- Text-and-image input only; no audio or video understanding, unlike some competing frontier multimodal models.
- Reachable only through xAI's own API and OpenRouter so far; none of the major managed-cloud platforms have added it yet.
Benchmarks
- apex swe: 56.4
- apex agents: 57.5
- aa briefcase: 1577
- deepswe v1 1: 65.9
- gdpval aa v2: 1753
- harvey lab vals: 15.8
- cursor bench v3 2: 69.9
- terminal bench v3 0: 26
- frontiercode v1 1 extended: 61.3
- artificial analysis intelligence index: 61
- artificial analysis price blended per m: 1.35
- artificial analysis speed tokens per sec: 67.6
Frequently Asked Questions
How much does Grok 4.6 cost per 1M tokens?
Grok 4.6's standard rate is $2 per million input tokens, $6 per million output, and $0.50 per million for cached input, all for prompts that stay under 200,000 tokens. Once a single request reaches that size, xAI bills the whole thing, not just the excess, at the long-context rate of $4 input, $1 cached, and $12 output per million. A faster variant is priced at roughly 2x standard, though xAI hasn't shared an exact rate or a separate model slug for it yet.
How does Grok 4.6 compare to GPT-5.6 Sol Max on benchmarks?
Grok 4.6 ties GPT-5.6 Sol Max for the top spot on the Artificial Analysis Intelligence Index, the closest head-to-head figure available since xAI hasn't published GPQA, AIME, or MMLU-Pro scores for either model. On xAI's own agentic suite, Grok 4.6 posts 57.5% on APEX-Agents and 61.3% on FrontierCode v1.1 Extended.
Is Grok 4.6 open source or proprietary?
Grok 4.6 is proprietary and API-only; xAI has not released its weights. Earlier Grok generations differ: Grok-1 shipped under Apache 2.0 and Grok-2's weights sit under the Grok 2 Community License, but neither predecessor's license extends to the 4.x line.
Does Grok 4.6 train on user data submitted through the API?
Enterprise customers can enable Zero Data Retention, which processes requests in real time without storing them; every API response includes an x-zero-data-retention header confirming whether it's active. Outside ZDR, xAI's terms allow it to create de-identified, aggregated data from API usage, and xAI states SOC 2 Type 2 compliance.
Who should use Grok 4.6, and who should avoid it?
Teams already running agent loops in Cursor or Grok Build get the most value, since both ship double included usage for the first week at Grok 4.5's price. Teams needing GPQA, AIME, or MMLU-Pro comparability, native audio or video input, or a managed-cloud endpoint on day one should look elsewhere until xAI publishes those or lands on Bedrock, Azure, or Vertex.
Top Alternatives
- Grok 4.5: Pick Grok 4.6 for the higher Artificial Analysis Intelligence Index score at the same price; pick Grok 4.5 only if you specifically need its Cursor co-training lineage.