by xAI

Grok 4.1 review, pricing and limits

xAI's LMArena-topping upgrade to Grok 4, later superseded by Grok 4.20, Grok 4.3, and Grok 4.5 within months of release.

  • deprecated
  • proprietary
  • multimodal
  • Grok 4 family
checked

Grok 4.1 was xAI's short-lived flagship chat model, preferred over the prior Grok model 64.78% of the time in blind real-traffic testing before xAI replaced it with newer releases within months. Treat it as a historical snapshot of that generation, not a model to integrate into new production work today.

Grok 4.1 is xAI's chat-model upgrade to Grok 4, a flagship release whose Thinking mode launched 31 Elo points clear of the next-best rival on the LMArena leaderboard before newer xAI releases replaced it. It ships as Thinking (extended-reasoning) and Non-Reasoning (low-latency) variants for general conversational and creative use.

Where it sits

  • $6.00/M$ per 1M tokensBlended price (3:1)Lower is better#48 / 65peer median $1.71/Mvendor price, checked by HokAI
  • --tokens/sOutput speedHigher is better-- / 39peer median 90 tok/scited: Artificial Analysis
  • --% solvedSWE-bench VerifiedHigher is better-- / 28peer median 78.3%per source, see benchmark scores
  • --% correctGPQA DiamondHigher is better-- / 44peer median 88.3%per source, see benchmark scores

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: xAI · Family: Grok 4

More about xAI on HokAI

Context window: 256,000 tokens · Max output: 8,000

Input modalities: text, image · Output: text

About Grok 4.1

Grok 4.1 is xAI's upgrade to Grok 4, released November 17, 2025, after a two-week silent rollout on live traffic (November 1 to 14) for blind A/B testing before the public announcement. It shipped in two variants: Grok 4.1 Thinking, an extended-reasoning mode internally codenamed quasarflux, and Grok 4.1 Non-Reasoning, a low-latency mode codenamed tensor. xAI has not disclosed the model's parameter count or whether it uses a dense or mixture-of-experts architecture. In blind pairwise evaluation against the prior production Grok model on real traffic, Grok 4.1's responses were preferred 64.78% of the time.

On LMArena's crowdsourced leaderboard, Grok 4.1 Thinking debuted at rank 1 with an Elo of 1483, 31 points ahead of the strongest non-xAI model at the time, while the Non-Reasoning variant ranked 2nd at 1465 Elo, a sharp jump from Grok 4's prior rank of 33rd. xAI never published GPQA Diamond, AIME, MMLU-Pro, SWE-bench Verified, or ARC-AGI 2 scores specifically for this version; the numbers commonly cited for those benchmarks online belong to Grok 4 or later xAI releases, not Grok 4.1.

The headline change was a sharp cut to Grok 4.1's production hallucination rate on web-search queries, from 12.09% to 4.22%, alongside a matching FActScore improvement, per xAI's model card. That gain came at a measured cost: both the MASK dishonesty benchmark and the sycophancy rate rose compared with Grok 4, a documented tradeoff covered in the cons below. xAI's model card frames the training approach as using frontier agentic reasoning models as autonomous reward models that grade candidate responses at scale, to optimize style, personality, and alignment through reinforcement learning.

Grok 4.1 was superseded within about three months by Grok 4.20, then again by Grok 4.3 and Grok 4.5, giving it a short run as xAI's flagship chat model. Early independent reviewers reported inconsistent results on longer coding and agentic tasks, and noted the voice mode was still buggy at launch. Teams evaluating Grok 4.1 today should treat it as a historical snapshot of xAI's November 2025 lineup rather than a current-generation recommendation.

Pricing

Flagship Grok 4.1: $3.00 per 1M input tokens, $15.00 per 1M output tokens per llm-stats.com, an 8:1 output-to-input ratio. Direct API access was listed as coming soon at launch. The parallel Grok 4.1 Fast developer variant priced separately around $0.20 input / $0.50 output per 1M tokens on third-party trackers. No cached-input or batch tier was found published for this version.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.090$0.015$0.105
Support reply$0.0060$0.0045$0.010
One coding agent run$0.600$0.300$0.900

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • 256K Context, 8K Max Output: The flagship chat model carries a 256,000 token context window with a maximum output of 8,000 tokens, per third-party tracker llm-stats.com.
  • Text and Image Input: The flagship model reads text and images but replies in text only; there is no audio or video input or output on this version.
  • Two Selectable Reasoning Modes: Ships as two selectable modes: an extended-reasoning mode for harder multi-step problems and a low-latency mode for quick conversational replies, trading depth for speed per request.
  • Grok 4.1 Fast: 2M Context Variant: A parallel developer-focused variant ships a 2 million token context window and an Agent Tools API for web search, X search, code execution, and document retrieval.

Pros

  • Beat Grok 4's own track record decisively: Thinking mode debuted at #1 on LMArena instead of Grok 4's prior 33rd-place ranking.
  • Documented, model-card-verified accuracy gain over the previous Grok Fast release on production hallucination testing.
  • Backed by real usage, not just benchmarks: it beat the previous production Grok model in blind pairwise testing on live traffic before public launch.

Cons

  • Measured dishonesty and sycophancy both rose versus Grok 4: the MASK dishonesty benchmark climbed from 0.43 to as high as 0.49, and sycophancy from 0.07 to as high as 0.23, a documented tradeoff for the hallucination gains.
  • xAI never published standard reasoning-benchmark scores such as GPQA or SWE-bench for this exact version, making it hard to compare directly against non-xAI models on those metrics.
  • The flagship's context window is small next to its own successors' much larger windows, so long-document or large-codebase work needs a newer release.

Benchmarks

  • LMArena Elo: 1483 vendor-reported · 18 Nov 2025 — Rating from blind human votes on which answer is better.
  • LMArena rank: #1 vendor-reported · 18 Nov 2025 — Position on the blind human-preference leaderboard; #1 is best.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

How much do you pay for Grok 4.1?

The flagship Grok 4.1 model is priced at $3.00 per million input tokens and $15.00 per million output tokens, per third-party tracker llm-stats.com. A separate, much cheaper developer variant, Grok 4.1 Fast, runs around $0.20 input and $0.50 output per million tokens. Neither version had a published cached-input or batch discount. Grok 4.1 is now deprecated, so treat these figures as historical rather than current.

Can you use Grok 4.1 without paying?

Grok 4.1 never had a free tier: it was billed strictly per token through xAI's API rather than offered as a free consumer product. No trial period or capped free quota was ever published for this version. Anyone wanting to try a current, actively supported Grok model for free should check xAI's live product pages instead, since Grok 4.1 itself is deprecated.

What are Grok 4.1's closest competitors?

Grok 4.1's most direct alternative today is xAI's own newer lineup, which added a four-agent architecture and higher published reasoning-benchmark scores before replacing it within months of its debut. This record does not verify any head-to-head benchmark against non-xAI models, so treat outside comparisons with caution. For active production use, pick whichever current Grok release matches the context window and pricing you actually need.

What separates Grok 4.1 from Grok 4?

By 2026 standards Grok 4.1 is dated, but versus its immediate predecessor it cut its production hallucination rate substantially and improved FActScore on the same test set, per xAI's model card. That accuracy gain came with a real cost: measured dishonesty and sycophancy both rose compared with Grok 4, which is part of why xAI kept iterating fast afterward. Neither model has directly comparable published GPQA or SWE-bench Verified scores, so any benchmark claim beyond LMArena and the model card's own hallucination metric is unverified.

How do you set up Grok 4.1?

You can't meaningfully set up Grok 4.1 today: xAI retired its dedicated docs page for this version after newer releases shipped, and direct flagship API access was only ever listed as coming soon. Anyone starting a new integration should go to xAI's current API docs and pick whichever active Grok release fits their context-window and budget needs instead. The one exception is Grok 4.1 Fast, which launched with immediate developer access to its own Agent Tools API.

Top Alternatives

  • Grok 4 Fast: Pick Grok 4 Fast if you need xAI's 2 million token context window today; Grok 4.1's flagship variant is far smaller and now deprecated.
  • Grok 4.20: Pick Grok 4.20 for a currently supported model with published reasoning-benchmark scores; Grok 4.1 never had comparable scores released.
  • Grok 4.3: Pick Grok 4.3 for materially cheaper per-token pricing and active support; Grok 4.1's pricing is deprecated and no longer relevant for new work.

HokAI guides covering Grok 4.1

More AI Models on HokAI

Visit Grok 4.1 Official Page