Grok 4.7 vs Claude Opus 5 vs GPT-6 Astra vs Gemini 3.8 Flash: What the Launch Numbers Don't Tell You
As of September 22, 2026, GPT-6 Astra leads the Artificial Analysis Intelligence Index at 53, ahead of Claude Opus 5 (51), Grok 4.7 (46) and Gemini 3.8 Flash (41). Claude Opus 5 leads coding specifically at 96.0% on SWE-bench Verified. Gemini 3.8 Flash is priced lowest at $0.75/$3.75 per million tokens.
The short version
Four frontier models launched within seven weeks, scoring 41-53 on the Artificial Analysis Intelligence Index. Gemini 3.8 Flash is the value pick, Claude Opus 5 wins unsupervised coding at 96% SWE-bench Verified, GPT-6 Astra costs the most for a narrow reasoning edge, and Grok 4.7's real advantage is distribution, not intelligence.
Four frontier models launched within seven weeks of each other: GPT-6 Astra on September 3, Gemini 3.8 Flash on September 2, and Grok 4.7 on September 21, all landing after Claude Opus 5's July 24 debut. Every vendor published an intelligence score at launch, and on the Artificial Analysis Intelligence Index the four now sit at 53, 51, 46 and 41.
That eight-point spread looks tight. The pricing behind it is not: Astra costs 13 times more per output token than the Gemini model for a lead of twelve points. That gap between the headline number and what a task actually costs is the story every launch post skipped.
Grok 4.7's Artificial Analysis scorecard, captured 22 Sept 2026. The 46 intelligence score and 38.8 tokens/sec speed sit below all three rivals covered here.
This guide is for a team choosing an API model for production work now, not for whoever wins next month's cycle. Short version: the Gemini model is the value pick for most workloads, Opus 5 is the one to reach for when a coding agent works unsupervised, Astra is worth its price only for tasks that need the extra reasoning headroom, and Grok 4.7's real launch-week advantage was reach, not intelligence.
The Verdict: Who Should Pick Which Model
In one pass, before the detail:
- High-volume, latency-sensitive work (drafting, extraction, classification): the Gemini model, for the price and the speed.
- An autonomous coding agent working unsupervised: Opus 5, for the SWE-bench Verified score.
- Research-grade reasoning or a contractual no-retention requirement: Astra, if the premium is worth it to you.
- A drop-in swap for a team already on Cursor or Grok Build: Grok, for the reach rather than the benchmarks.
Pick Gemini 3.8 Flash if you are running high-volume, latency-sensitive work: support drafts, data extraction, first-pass classification. At $0.75 input and $3.75 output per million tokens, generating at 328.8 tokens per second, it costs a fraction of the alternatives and still lands a score of 41.
Pick Claude Opus 5 if the job is an autonomous coding agent that has to open pull requests without a human checking every step. It posted 96.0% on SWE-bench Verified in Anthropic's own system card, the strongest coding figure of the four, at a mid-pack $5/$25 per million tokens.
Pick GPT-6 Astra only when the task genuinely needs the extra reasoning headroom over Opus 5: novel scientific reasoning, long agentic chains where small errors compound, or work where a contractually guaranteed no-retention mode matters. At $10/$50 per million tokens it costs the most of the four, and the premium buys a real but narrow edge.
Grok 4.7 is the pick for reach rather than raw capability. It shipped day one inside Cursor and Grok Build alongside the standard API, at $2/$6 per million tokens, according to xAI's own launch materials as covered by Kingy AI. If your team already lives inside those tools, that convenience is real. Choosing an API model from a blank slate, its 46 score and the coding numbers below put it behind the other three.
Opus 5's scorecard on the same tracker, pulled the same day. The 1M token window is double Grok's 500K.
What Each Model Actually Costs
Vendors quote input and output rates separately, and the gap between them is where buyers get surprised. Figures below are per million tokens, current as of the dates shown.
| Model | Input /1M | Output /1M | Context window | Released |
|---|---|---|---|---|
| Gemini 3.8 Flash | $0.75 | $3.75 | 1M tokens | Sept 2, 2026 |
| Grok 4.7 | $2.00 | $6.00 | 500K tokens | Sept 21, 2026 |
| Claude Opus 5 | $5.00 | $25.00 | 1M tokens | July 24, 2026 |
| GPT-6 Astra | $10.00 | $50.00 | 1M tokens | Sept 3, 2026 |
Source: Artificial Analysis model pages, accessed September 22, 2026.
Grok's context window is genuinely the odd one out at 500,000 tokens against roughly 1 million for the other three. Its pricing also doubles above a 200,000-token prompt, according to CellCog's coverage of the launch, so a long-context workload on this specific model costs more than the headline rate suggests.
A worked example makes the output-token gap concrete. A support team processing 1,000 tickets a day, averaging 500 input tokens and 200 output tokens per ticket, spends roughly this much per day at each vendor's rates:
| Model | Daily cost (1,000 tickets) | Monthly cost |
|---|---|---|
| Gemini 3.8 Flash | $1.13 | ~$34 |
| Grok 4.7 | $2.20 | ~$66 |
| Claude Opus 5 | $7.50 | ~$225 |
| GPT-6 Astra | $15.00 | ~$450 |
That is a $34 line item against a $450 one for identical ticket volume. The gap only closes if the cheaper answers need enough human correction to erase the savings, a real risk beyond routine drafting.
The Intelligence Index Gap Is Smaller Than the Price Gap Suggests
Artificial Analysis measures all four on the same ten-benchmark composite. As of September 22, 2026: GPT-6 Astra scores 53, Claude Opus 5 scores 51, Grok 4.7 scores 46, and Gemini 3.8 Flash scores 41.
Look at what that buys. A two-point edge for the top scorer costs exactly double the price per output token. The cheapest of the four captures 77% of that top score at 7.5% of its output cost. Grok sits in the middle on price and near the bottom on the composite, the least comfortable spot on the chart: not the cheapest, not the smartest.
Anthropic's own smaller model makes the same point from a different angle. In the same benchmark round covered by The Decoder, Claude Fable 5.1, a cheaper release than Opus 5, scored 55% on Terminal-Bench 4.0, beating Grok's 26% on the identical test. A company's second-tier model beating a same-week frontier launch on a coding benchmark is exactly what a single composite score hides.
Where Coding Ability Actually Diverges
The composite score is a blend. For engineering teams, the coding-specific numbers matter more, and here the benchmark version matters as much as the result. Terminal-Bench 4.0 is a newer, harder revision than 2.1, so scores across the two do not compare directly.
| Model | Benchmark | Score |
|---|---|---|
| GPT-6 Astra | Terminal-Bench 4.0 | 57.7% |
| Claude Fable 5.1 (not Opus 5) | Terminal-Bench 4.0 | 55% |
| Grok 4.7 | Terminal-Bench 4.0 | 26% |
| Claude Opus 5 | SWE-bench Verified | 96.0% |
Source: The Decoder's benchmark coverage and Anthropic's system card via SeaWork, both accessed September 22, 2026. Opus 5 was not run against version 4.0 in the sources checked for this guide, so it is absent from that column. SWE-bench Verified and Terminal-Bench measure related but different things, patch accuracy against a fixed test suite versus open-ended terminal task completion, so read the two rows as complementary rather than one ranking.
The practical read: if your coding workload looks like SWE-bench, resolving a defined issue against a known repo, that 96.0% figure is the strongest signal available right now. If it looks like Terminal-Bench, open-ended terminal and tool use, the 57.7% score currently leads the field on that exact test.
Where Grok 4.7 Wins: Reach, Not Intelligence
Grok's real story at launch was not the benchmark table. It shipped with same-day access through Cursor, Grok Build and the standard API, according to launch coverage from Kingy AI, matching what the vendor itself published. That is a faster path into an existing coding workflow than waiting on a separate integration.
One claimed advantage does not hold up under a second source. Musk said SpaceX engineering data was folded into training, but a CellCog review of the actual launch post found no mention of that dataset in the vendor's own materials. This guide treats the claim as unconfirmed rather than repeating it as fact.
What is confirmed is pricing that matched the prior release exactly and a same-day rollout into tools developers already use. For a team already paying for those seats and needing a drop-in swap with no new procurement step, that convenience has real value. Our earlier coverage of the coding distribution deal goes deeper on what changed there. Just do not confuse fast reach with strong benchmarks.
Where Claude Opus 5 Wins: Coding You Can Trust Unsupervised
SWE-bench Verified is built to resist gaming: a fixed set of real GitHub issues with a known correct patch, penalizing a model that claims success without resolving the issue. Opus 5's 96.0% there, reported in Anthropic's own system card, is the strongest signal in this comparison for one specific job: an autonomous agent that opens a pull request without a human reviewing every intermediate step.
Its million-token context window also matches the two priciest models here rather than trailing Grok's smaller one, which matters for an agent that needs a whole repository in view rather than a retrieved slice of it. For how this class of model stacks up against a dedicated coding-agent product, our Cursor Composer vs Claude Code piece is the closer look, and HokAI's coding-assistant category covers the wider field of products built on top of models like this one.
Where GPT-6 Astra Wins: The Extra Reasoning Headroom
Astra's composite score of 53 is the highest of the four, and it holds the best published Terminal-Bench 4.0 result at 57.7%. OpenAI also offers eligible enterprise customers a mode that deletes inputs and outputs the moment a request finishes, with no staff review, rather than the 30-day default most of the field uses.
The honest read is that this is a narrow lead bought at a steep premium. A few points of composite score and a coding-benchmark edge do not obviously justify double the output price for every workload. It earns that cost on tasks where the extra headroom decides the outcome: research-grade reasoning chains, or agentic runs long enough that a small per-step error rate compounds into a wrong final answer.
Where Gemini 3.8 Flash Wins: Price and Speed, Not a Close Second
Gemini 3.8 Flash generates at 328.8 tokens per second, more than eight times Grok 4.7's 38.8 and nearly five times GPT-6 Astra's 66.6, per the same benchmark provider. Paired with its $0.75/$3.75 pricing, reach for it when the job is high-volume and latency-sensitive rather than reasoning-bound: drafting, extraction, summarization, anything run at scale where a score of 41 is enough headroom.
For scale, Google's own previous release, Gemini 3.6 Flash, scored 34 on the same index at launch and has since been deprecated in favor of newer Flash releases; the six-week jump to 41 shows how fast this specific product line moves. At the very bottom of the price range, Mistral Large 3 from Mistral AI undercuts even the Gemini model at $0.50/$1.50 per million tokens, though Mistral has not published a comparable composite score for it, so that comparison stops at price.
The Fine Print: These Are the Top-Effort Scores
Every number above is the highest reasoning-effort configuration each vendor publishes, not necessarily what a default API call returns. The tracker labels its test runs accordingly: Grok's entry is marked "xhigh," Astra's is marked "max," and the Gemini model's is marked "high." Opus 5's listing specifies "Adaptive Reasoning, Max Effort," and the tracker's own test run against that setting needed 140 million output tokens to complete, well above the roughly 94 million median across comparable runs.
That matters for two reasons. First, a lower effort setting on any of the four will typically score below the numbers in this guide and cost less per response, since more reasoning tokens are what the higher tiers spend to earn their extra points. Second, it means the price comparisons above are apples to apples: all four are quoted at the configuration that produced the composite score being compared, not a cheaper default that would understate the real cost of matching that score.
For most production workloads, that top tier is overkill. A support-ticket classifier does not need the priciest model's maximum setting any more than a spreadsheet needs a supercomputer. Test each vendor's default and low-effort tiers against your own task before assuming the number in this guide is the price you will actually pay.
How Each Model Handles Your Data
All four vendors say API inputs are not used to train their models by default, but the retention terms behind that promise differ in ways that matter for regulated workloads.
| Model | Default retention | Enterprise no-retention option | Data residency |
|---|---|---|---|
| Claude Opus 5 | Standard API policy | Available to eligible accounts | US, EU |
| GPT-6 Astra | Not used for training | Deletes data on request completion, no staff review | Not published |
| Grok 4.7 | Encrypted, 30 days, then deleted | Available, but disables stateful files/collections/batch | US endpoint (+10% price) |
| Gemini 3.8 Flash | Paid tier: not used for training | N/A; free tier may train and be reviewed | Not published |
Anthropic classifies Opus 5 under the EU AI Act as a general-purpose system with systemic-risk obligations, and offers both US and EU residency. Google's exception matters for the free tier specifically: users in the EEA, Switzerland and the UK get the paid-tier terms applied even on unpaid usage.
If your workload cannot tolerate any server-side retention, the two priciest models here currently offer the cleanest enterprise paths. Grok's zero-retention setting costs you the stateful features that make an agent workflow convenient in the first place.
The Turn: Maybe the Composite Score Doesn't Matter Here
The strongest objection to everything above is that a ten-benchmark blend was never built to predict performance on your specific task. A model scoring five points lower on a broad average can still be the better choice if its individual strengths line up with what you are building, and a five-point average gap can come from one or two tests that have nothing to do with your use case.
That objection is fair, and it is exactly why this guide breaks the comparison into coding, price, speed and data handling separately rather than stopping at one number. The composite is a useful first filter for narrowing four models to two worth testing. It is a poor substitute for running your own eval set against the finalists.
There is a second, quieter problem with any single tracker's score: different evaluation services do not always agree with each other, because they weight benchmarks differently and update their suites on their own schedule. A model that leads one composite by two points might trail a rival composite by two points the same week, without either service being wrong.
Treat any single number here as a snapshot from one respected source, not a settled fact about which model is smarter in some absolute sense. The benchmark tables further up this guide, sourced independently and dated, are the more durable evidence. A single composite score is only the quick filter that gets you to those tables, not the final word.
Switching Cost: What Moving Between These Four Actually Takes
All four sit behind a standard REST API with OpenAI-compatible or vendor-native SDKs, so the code-level switching cost is genuinely low for a simple completion call. It rises fast once a workflow depends on vendor-specific features.
Moving off Grok's zero-retention setting means giving up its stateful file, collection and batch features, which can mean rewriting an agent's state-management layer. Moving between the other two vendor families means re-tuning prompts: they respond differently to system-prompt structure and tool-call formatting, and a prompt tuned for one rarely transfers cleanly to the other.
The context-window gap matters here too. An agent built around a million-token window, feeding it a whole repository or document set at once, cannot drop into the smallest context window here without adding a retrieval or chunking step first. That is an afternoon of engineering for a small workflow and a real re-architecture for a large one. If you are choosing between models under a specific constraint, open weights or the lowest possible latency among them, HokAI's model recommender walks through that decision by constraint rather than by brand name.
What Would Change This Verdict
This comparison is a snapshot of four launches inside a seven-week window, and the numbers behind it move. If the priciest model's rate drops toward the mid-pack price, as prior generations from the same vendor have within months of launch, its lead in reasoning stops being a premium purchase and becomes the default pick outright.
If a later published result on Terminal-Bench 4.0 for the Anthropic model lands below the current 57.7% leader, the coding verdict above needs revisiting. And if Grok closes that same gap the way the Cursor coding integration covered earlier already closed the distribution gap, its price-to-capability ratio gets a lot more interesting than it is today.
For now, the fastest way to see how these four, and everything else HokAI tracks, stack up on the metric you actually care about is the live model leaderboard, which updates as new scores land rather than staying frozen at launch-day numbers the way this guide necessarily is. If you are unsure which axis matters most for your project, HokAI's Smart Match asks about your actual workload and narrows the field from there instead of asking you to read four spec sheets first.
Frequently asked questions
Which model scores highest on the Artificial Analysis Intelligence Index?
GPT-6 Astra scores 53, the highest of the four, as of September 22, 2026. Claude Opus 5 follows at 51, Grok 4.7 at 46, and Gemini 3.8 Flash at 41. The gap between the top and bottom score is smaller than the price gap between the same two models.
Is Grok 4.7 good at coding?
It trails the other three on the coding benchmarks checked for this guide: 26% on Terminal-Bench 4.0, compared with GPT-6 Astra's 57.7% and even Anthropic's smaller Claude Fable 5.1 at 55% on the same test. Its real strength at launch was distribution, shipping instantly inside Cursor and rolling out across GitHub Copilot, not raw coding capability.
Which of these four is cheapest to run?
Gemini 3.8 Flash, at $0.75 per million input tokens and $3.75 per million output tokens. That is roughly 13 times cheaper on output tokens than GPT-6 Astra, which is priced at $10/$50 per million.
Do Grok 4.7, Claude Opus 5, GPT-6 Astra and Gemini 3.8 Flash train on my data?
None of the four train on API inputs by default, according to each vendor's published terms. Grok 4.7 retains encrypted API traffic for 30 days for abuse monitoring before deleting it; GPT-6 Astra and Claude Opus 5 both offer Zero Data Retention to eligible enterprise customers; Gemini 3.8 Flash's paid tier does not train on inputs, though its free tier may.
Which model should I actually use for an autonomous coding agent?
Claude Opus 5, based on its 96.0% SWE-bench Verified score reported in Anthropic's own system card, the strongest published result among the four for resolving real GitHub issues without a human checking every step. Its million-token context window also matters for agents that need a whole repository in view rather than a retrieved slice of it.
Covered in this guide
- Grok 4.7: xAI's September 2026 flagship model, built for coding and agentic work with a 500K token context window and a 46.3% CursorBench score.
- Claude Opus 5: Anthropic's July 2026 flagship LLM, with a 1M token context window by default and a new xhigh reasoning-effort mode for long agentic runs.
- Gemini 3.8 Flash: Google DeepMind shipped Gemini 3.8 Flash on September 2, 2026, a multimodal Gemini 3 model tuned for long-horizon coding and agentic enterprise work.
- GPT-6 Astra: OpenAI's flagship model, launched September 2026 as the first ever rated at the Preparedness Framework's Critical cybersecurity level.
- Anthropic: Anthropic, founded 2021 by 7 ex-OpenAI researchers, builds Claude and was valued near $965B after its May 2026 Series H round.
- Cursor: Cursor is an AI code editor built on VS Code, widely adopted across large enterprise engineering teams, with Agent Mode, Tab completion, and Cloud Agents.
- Gemini 3.6 Flash: Released July 21, 2026, Gemini 3.6 Flash is Google DeepMind's Flash-tier model built for agentic coding, computer use and long-context work.
- Mistral AI: Mistral AI, founded in April 2023 in Paris by three ex-Meta researchers, builds Mistral, Mixtral, and Le Chat and raised $1.47B including $830M debt (Mar 2026).
- Mistral Large 3: Mistral Large 3 (Dec 2025) is a 675B MoE model (41B active) with 256K context and Apache 2.0 open weights.
- OpenAI: OpenAI builds the GPT-5.6 model family (Sol, Terra, Luna), o3, ChatGPT (900M+ weekly users), and the OpenAI API. Closed a $122B round at an $852B valuation in March 2026, the largest private funding round in history.
- xAI: Elon Musk's AI company (roughly 4,000 to 4,900 employees) builds the Grok models and merged into SpaceX, taking the combined business public on Nasdaq as the largest IPO on record.
Sources
Still deciding?
This guide covers a handful of options. Smart Match checks every listing in the directory against how you actually work and what you can spend, then hands you the shortlist and the reason behind each pick.
Start Smart MatchRelated guides
- AI Development Services in 2026: Which Layer You Actually NeedBuyer's guideHow to pick, across a category
- Best AI Chatbots in 2026: Pick by the Job, Not the LeaderboardBuyer's guideHow to pick, across a category
- Best AI Coding Assistants in 2026: Pick the Job, Not the BrandBuyer's guideHow to pick, across a category
- Best AI Companies in 2026: Who Is Actually LeadingBuyer's guideHow to pick, across a category
- Best AI for Writing a Business Plan in 2026: Tested Picks and a VerdictBuyer's guideHow to pick, across a category
- Best AI for Coding Questions Free in 2026: 8 Real Options, ComparedBuyer's guideHow to pick, across a category