by xAI

Grok 4.7 review, pricing and limits

xAI's flagship model for agentic coding, engineering, and knowledge work, and the default model behind Cursor and Grok Build.

  • ga
  • proprietary
  • multimodal
  • Grok 4 family
checked

Grok 4.7 launched September 21, 2026 with a 500,000 token context window and four reasoning effort levels from low to xhigh. It extends Grok 4.6 with a larger base model and a longer reinforcement learning run, and is the default model behind Cursor and Grok Build's coding agents.

Grok 4.7 is xAI's flagship large language model, released September 21, 2026, scoring 46.3% on CursorBench 4.0 at xhigh effort and 71.0% on DeepSWE v1.1. It has a 500,000 token context window, accepts text and image input, and is the default model in Cursor and Grok Build.

Where it sits

  • $3.00/M$ per 1M tokensBlended price (3:1)Lower is better#37 / 64peer median $1.70/Mvendor price, checked by HokAI
  • 39 tok/stokens/sOutput speedHigher is better#35 / 39peer median 90 tok/scited: Artificial Analysis
  • --% solvedSWE-bench VerifiedHigher is better-- / 28peer median 78.3%per source, see benchmark scores
  • --% correctGPQA DiamondHigher is better-- / 44peer median 88.3%per source, see benchmark scores

Priced around the middle of the 64 GA models with a published price (rank 37), rank 35 of 39 on output speed as cited from Artificial Analysis, and one of 22 that document a zero-data-retention option.

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: xAI · Family: Grok 4

More about xAI on HokAI

Context window: 500,000 tokens

Input modalities: text, image, tool-calls · Output: text, tool-calls

About Grok 4.7

Grok 4.7 is xAI's (doing business as SpaceXAI) latest large language model, released September 21, 2026. It extends Grok 4.6, released August 12, 2026, with a new, larger base model trained with a longer reinforcement learning run on a harder mix of tasks weighted toward problems that take multiple hours to complete. xAI describes it as the company's most capable model to date, built primarily for coding, agentic tasks, and knowledge work. The model received supplemental training on anonymized Cursor workflow data specifically to sharpen coding and agentic performance, and was trained to natively understand the Grok Bot agent runtime for conversational and general knowledge work.

On CursorBench 4.0, a benchmark of long-horizon coding tasks pulled from real Cursor sessions, Grok 4.7 scores 46.3% at xhigh reasoning effort and 43.9% at high, ahead of Grok 4.6's 40.4% but behind Anthropic's Fable 5.1 at 51.8%. On DeepSWE v1.1, a 113-task software engineering benchmark graded by an independent verifier container, it scores 71.0% at high effort, ahead of Grok 4.6's 65.2% but behind GPT-5.6 Sol's 72.7%. On the Artificial Analysis Intelligence Index (v4.3.2, a 10-benchmark composite), the xhigh configuration scores 46, generating output at 38.8 tokens per second in Artificial Analysis's independent testing. xAI also reports EEBench (electrical and chip design) at 66.0%, CADGenBench at 44.4%, and SWE-Marathon v1.1 at 46.0%, each near the top of the tested peer set of Grok 4.6, Fable 5.1, GPT-5.6 Sol, and Opus 5.

Grok 4.7 has a 500,000 token context window and, per xAI's release notes, no separate output token limit is enforced beyond that shared budget. Requests that exceed 200,000 tokens of context are billed at a higher rate on the standard API. The model always returns encrypted reasoning content (reasoning.encrypted_content) regardless of whether it was requested, which matters for anyone building multi-turn agentic tool-calling loops that carry reasoning state between calls.

Grok 4.7 accepts text and image input and produces text-only output; audio and video are handled by separate grok-imagine-voice and grok-imagine-video models rather than by Grok 4.7 itself. It supports function calling, structured outputs, reasoning at four effort levels (low, medium, high, xhigh, with high as the default), and a set of built-in agentic tools described under Key Features below.

Standard pricing matches Grok 4.6's rate card, with a premium tier once a request's context grows past 200,000 tokens (see the pricing FAQ for exact figures). A faster, pricier variant exists for Cursor and Grok Build users only, not the public API. The Batch API, supported on some other xAI models, is not available for Grok 4.7. A US-only regional endpoint (us.api.x.ai/v1) carries a premium over the global rate and excludes image, video, and voice APIs.

Beyond the direct API at console.x.ai, Grok 4.7 is the default model in Grok Build (xAI's terminal coding agent) and in Cursor on every plan tier, and is the default model in the Grok add-ins for Microsoft Word, PowerPoint, and Excel. It is also reachable through Google Cloud Vertex AI (as a partner model in Model Garden, model ID xai/grok-4.7, via an OpenAI-compatible interface) and Microsoft Azure AI Foundry, plus model gateways including OpenRouter, Vercel, Cloudflare, Snowflake, and Databricks Mosaic. Direct API rate limits are 150 requests per second and 50 million tokens per minute per region, across us-east-1, us-west-2, and us-central-1.

Training combined supervised fine-tuning and reinforcement learning on human and synthetic reward signals, with model-based checks screening out problematic SFT traces generated by Grok 4.6. xAI's published model card reports a layered, defense-in-depth safeguard stack: standard jailbreak compliance of 0.01%, StrongREJECT compliance of 2.0%, and long-horizon multi-turn jailbreak compliance of 0.65%. On CBRN-specific evaluations, bio refusal recall is 100%, chemical refusal recall is 99.9%, and radiological/nuclear refusal accuracy (FORTRESS-RN) is 97.9%; child-safety compliance is 0.0%. On HackerBench v0.3, xAI's internal red-team cyber suite, harmful/dual-use compliance at high effort is 3.31%. Nine external partners (Abundant AI, Atopile, Cathedral, Datacurve, Harbor, LatchBio, Mecado, Proximal Labs, and Vals AI) ran independent evaluations for the model card, and cyber capabilities were separately corroborated by third-party evaluators under xAI's Frontier Artificial Intelligence Framework, published June 30, 2026.

Teams building agentic coding tools on top of Cursor or Grok Build get native, first-class support and the model's strongest measured results, across CursorBench, DeepSWE, and Terminal-Bench 4.0. Electrical and CAD engineering teams have a genuine, benchmarked use case in EEBench and CADGenBench, where Grok 4.7 leads or nearly leads every tested peer. Teams whose workloads center on ultra-long-horizon reasoning or ground-up systems work may get better results elsewhere, since Fable 5.1 leads clearly on FrontierSWE V2 (56.3% versus 29.0%) and Terminal-Bench 4.0 (57.9% versus 38.0%). Voice or video-first products need a separate model, since Grok 4.7 has no native audio or video I/O, and teams that depend on asynchronous batch processing will need a workaround since the Batch API is unavailable for this model.

Grok 4.7 has a pretraining data cutoff of June 2026 and uses supplemental training data generated as late as August 2026, drawn from public web data, internally generated data, and data xAI has secured rights to. By default, API requests and responses are stored encrypted at rest for 30 days for abuse monitoring and then deleted; xAI states it does not train on API inputs or outputs without explicit permission. Enterprise customers can enable Zero Data Retention, which disables the stateful Responses API, Files, Collections, and Batch API in exchange for no server-side storage at all. xAI states it is SOC 2 Type 2 compliant and offers HIPAA Business Associate Agreements on request; its own FAQ stops short of confirming GDPR compliance directly, referring customers with a signed NDA to its Trust Center instead.

Pricing

Standard rate is $2.00 input / $0.50 cached input / $6.00 output per 1M tokens, matching Grok 4.6. Requests beyond 200,000 tokens of context bill at $4.00 / $1.00 / $12.00 per 1M tokens. A Grok 4.7 Fast variant runs at double the token price for double the output speed, but is only reachable through Cursor and Grok Build, not the public API. The Batch API is not supported for this model. The US-only regional endpoint (us.api.x.ai/v1) is priced 10% above the global rate.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.060$0.0060$0.066
Support reply$0.0040$0.0018$0.0058
One coding agent run$0.400$0.120$0.520

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • Four reasoning effort levels: Low, medium, high, and xhigh let you trade cost and latency for accuracy on a per-request basis, with high as the default.
  • 500K token context window: Shares a single 500,000 token budget across input and output, with no separate output token cap.
  • Native Cursor and Grok Build integration: Default model in Cursor on every plan tier and in xAI's own Grok Build terminal coding agent.
  • Multi-cloud availability: Reachable through xAI's own API, Google Cloud Vertex AI Model Garden, and Microsoft Azure AI Foundry.
  • Server-side agentic tools: Built-in web search, X search, code execution, image generation, RAG collections search, and remote MCP tool support.

Pros

  • Scores 71.0% on DeepSWE v1.1 (high effort) and 66.0% on EEBench (xhigh), leading or near-leading its tested peer set on coding and electrical-engineering benchmarks.
  • Default model in Cursor across every plan tier and in Grok Build, giving it broad, frictionless distribution into coding workflows.
  • Strong measured jailbreak resistance: 0.01% standard jailbreak compliance and 100% bio / 99.9% chem CBRN refusal recall per its own model card.

Cons

  • No native audio or video input or output, unlike some multimodal competitors.
  • No asynchronous Batch API access, unlike some other xAI models.
  • Trails Anthropic's Fable 5.1 by a wide margin on longer-horizon benchmarks such as FrontierSWE V2 and Terminal-Bench.

Benchmarks

  • AA Intelligence Index: 46 cited: Artificial Analysis · 22 Sep 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
  • AA blended price: $1.35/M cited: Artificial Analysis · 22 Sep 2026 — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.
  • Output speed: 39 tok/s cited: Artificial Analysis · 22 Sep 2026 — Median tokens written per second as measured by Artificial Analysis.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

What does Grok 4.7 cost per million tokens?

Grok 4.7 costs $2.00 per 1M input tokens and $6.00 per 1M output tokens on xAI's standard API, with cached input at $0.50 per 1M tokens, the same rate as Grok 4.6. Requests beyond 200,000 tokens of context are billed at double the standard rate. A separate Grok 4.7 Fast variant, available only through Cursor and Grok Build, runs at twice the price for twice the output speed.

How does Grok 4.7 compare to GPT-5.6 Sol and Fable 5.1 on coding benchmarks?

On CursorBench 4.0, Grok 4.7 scores 46.3% against Fable 5.1's 51.8% and GPT-5.6 Sol's 41.7%, ahead of GPT-5.6 Sol but behind Fable 5.1. On DeepSWE v1.1, GPT-5.6 Sol edges it out at 72.7% versus Grok 4.7's 71.0%. Fable 5.1 leads more clearly on longer-horizon work, scoring 56.3% on FrontierSWE V2 versus Grok 4.7's 29.0%.

Is Grok 4.7 open source?

No. Grok 4.7 is a proprietary model available through xAI's own API and hosted surfaces (Cursor, Grok Build, Office add-ins), or through partner platforms like Google Cloud Vertex AI and Microsoft Azure AI Foundry. xAI has not disclosed the model's parameter count or released its weights.

Does xAI train Grok 4.7 on customer API data?

By default, no. xAI stores API requests and responses encrypted at rest for 30 days for abuse monitoring, then deletes them automatically, without using that data for training. Enterprise teams can enable Zero Data Retention for no server-side storage at all, though doing so disables features like the stateful Responses API, Files, Collections, and the Batch API.

Who should use Grok 4.7, and who should look elsewhere?

Teams building agentic coding tools in Cursor or Grok Build, or doing electrical and CAD engineering work, get Grok 4.7's strongest measured results. Teams needing native voice or video input, large asynchronous batch jobs, or the top score on ultra-long-horizon reasoning work are better served by a dedicated audio/video model, a Batch-API-supporting alternative, or Fable 5.1 respectively.

HokAI guides covering Grok 4.7

More AI Models on HokAI

Visit Grok 4.7 Official Page