Claude Mythos 5.1 review, pricing and limits

Anthropic's top-tier model for defensive cybersecurity work and biosecurity research, gated behind Project Glasswing's vetting process.

  • ga
  • proprietary
  • multimodal
  • Claude Mythos family
checked

Released September 1, 2026, Claude Mythos 5.1 pairs a 1 million token context window with 128,000 token output limits, invite-only for vetted cybersecurity defenders and biology researchers under Project Glasswing. It replaces ad hoc manual vulnerability triage for security teams that already qualify for Anthropic's trusted access programs, not a general-purpose assistant for public API customers.

Claude Mythos 5.1 is Anthropic's Mythos-class large language model, scoring 90% on ARC-AGI-2 at maximum reasoning effort as independently verified by the ARC Prize foundation. It shares Claude Fable 5.1's weights but relaxes cybersecurity and biology safety classifiers for vetted defenders and researchers, unlike Anthropic's publicly available API models.

Where it sits

  • $20.00/M$ per 1M tokensBlended price (3:1)Lower is better#65 / 69peer median $1.71/Mvendor price, checked by HokAI

Pricier than 97% of the 69 GA models with a published price, and one of 26 that document a zero-data-retention option.

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: Anthropic · Family: Claude Mythos

More about Anthropic on HokAI

Context window: 1,000,000 tokens · Max output: 128,000

Input modalities: text, image · Output: text

About Claude Mythos 5.1

Anthropic released Claude Mythos 5.1 on September 1, 2026 as a restricted-access large language model. It shares its underlying weights with Claude Fable 5.1: the two are the same model deployed under different safeguard configurations. Fable 5.1 is generally available with Anthropic's standard cybersecurity and biology classifiers active, while Mythos 5.1 relaxes those classifiers for vetted researchers and defenders enrolled in Anthropic's Cyber Verification Program or Life Sciences Verification Program, together known as Project Glasswing. Anthropic positions it as its most capable model for cybersecurity defense and life sciences research, sitting alongside Claude Opus 5.5 and Claude Sonnet 5 in the Claude 5 lineup but reserved for gated, invitation-only access under Project Glasswing rather than the standard API.

On Anthropic's own published evaluations, Fable 5.1 (and by extension Mythos 5.1, which shares its weights) scores 1,853 on GDPval-AA v2, a real-world knowledge-work benchmark, up from 1,723 for the previous Fable 5 generation. It scores 65.0% on Humanity's Last Exam when allowed to use tools, and 60.9% without tools. Independent testing from the ARC Prize foundation recorded 97.5% on ARC-AGI-1 and 90.0% on ARC-AGI-2 at maximum reasoning effort, each run costing $1.40 and $4.49 per task respectively. Third-party leaderboards including BenchLM and CodingFleet recorded 81.2% on SWE-bench Pro as of September 2026. Anthropic has not published SWE-bench Verified, GPQA Diamond, AIME or MMLU-Pro scores for this release, so those figures are left blank rather than estimated.

The model has a 1 million token context window and a 128,000 token maximum output on the synchronous Messages API, extendable to 300,000 output tokens on the asynchronous Message Batches API with a beta header, a limit shared with several other Claude models. Anthropic lists its reliable knowledge and training data cutoff as June 2026. Thinking is adaptive and always on with a default effort of high, and Anthropic's own comparison table marks its latency as slower than the rest of the Claude 5 lineup. Input is text and images; output is text only, the same split Anthropic ships across Fable 5.1, Opus 5.5, Sonnet 5 and Haiku 4.5.

Mythos 5.1's defining difference from Fable 5.1 is what its safeguards allow: Anthropic's cybersecurity classifiers, which block penetration testing, exploit generation and binary-based vulnerability scanning by default, are tuned to let Cyber Verification Program participants identify vulnerabilities directly in their own source code for defensive work, and its biology safeguards intervene on benign requests 85% less often than the Fable 5 generation, according to Anthropic. Pricing follows the same per-token structure as the rest of the Claude 5 lineup, and Mythos 5.1 and Fable 5.1 are priced identically since they are the same deployment of one underlying model; exact current rates are in the pricing FAQ below. There is no free tier, since access itself is gated by program enrollment rather than payment.

Mythos 5.1 is available through the Claude API, Amazon Bedrock, Google Cloud Vertex AI and Microsoft Foundry, under the model ID claude-mythos-5-1 (anthropic.claude-mythos-5-1 on Bedrock). Unlike the rest of the Claude 5 lineup, it cannot be reached with a standard API key: every request needs prior Project Glasswing approval, and that approval is currently open only to organizations headquartered in the US, with international expansion planned (see the access FAQ below for how to apply). Anthropic tested Fable 5.1's cybersecurity safeguards internally, with two external organizations, and with Gray Swan, and reported no critical-severity jailbreak, consistent with Fable 5 and Opus 5; the joint system card documents expert red-teaming of the model's operational security and biology research capabilities.

Even inside the Cyber Verification Program, Mythos 5.1 draws a firm line at offensive tooling: no functioning exploits, no live intrusion attempts, no automated bug-hunting in compiled binaries; the new allowance is limited to identifying vulnerabilities directly in source code for defensive purposes. Ambiguous dual-use biology or chemistry questions route to Anthropic's Opus models instead of being answered directly, and new API accounts on Mythos 5.1 and Fable 5.1 cannot manually edit context, an anti-distillation safeguard Anthropic added for this release. Mythos 5.1 fits security teams already vetted for the Cyber Verification Program who need to triage vulnerabilities in their own source code, and life-science researchers in the Life Sciences Verification Program running literature synthesis, protein design or genome-wide analysis. General application development is better served elsewhere: Claude Sonnet 5 and Claude Opus 5.5 cost far less per token, need no vetting, and cover ordinary coding and reasoning work without waiting on program approval.

Anthropic's commercial products, including the Claude API that Mythos 5.1 runs on, carry SOC 2 Type I and Type II, ISO 27001:2022 and ISO/IEC 42001:2023 certification and support HIPAA business associate agreements, per Anthropic's privacy center. API inputs and outputs are not used to train Anthropic's models by default; Mythos 5.1 adds a 30-day retention window specifically for safety monitoring of Project Glasswing usage on top of Anthropic's standard commercial retention policy, and Anthropic's Enterprise Frontier Safeguards option allows zero-data-retention deployment on a customer's own infrastructure.

Fable 5.1 and Mythos 5.1 replaced Fable 5 and Claude Mythos 5 on September 1, 2026, with export controls on the prior Mythos generation having been lifted July 1, 2026 after US government approval. Anthropic's own figures show Terminal-Bench-Science 0.1 more than doubling from 24.7% to 52.6%, alongside a steep cut to cache-read pricing versus the previous generation. The disclosed tradeoff is token usage: at maximum reasoning effort, Fable 5.1 and Mythos 5.1 use about 1.7 times the output tokens of Fable 5 for the same task, and cost about 20% more per task despite the cheaper cache reads.

Pricing

Input costs $10 per million tokens and output $50 per million, the same rate as Claude Fable 5.1 since both are priced identically. A 5-minute prompt cache write costs $12.50 per million tokens, a 1-hour cache write $20 per million, and cache reads $0.25 per million, a 75% cut from the previous generation. The asynchronous Batch API gives a 50% discount on both input and output. There is no free tier; access is invitation-only through Project Glasswing.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.300$0.050$0.350
Support reply$0.020$0.015$0.035
One coding agent run$2.00$1.00$3.00

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • 1 Million Token Context: Processes up to 1 million tokens of input, the same window Anthropic ships across the whole Claude 5 lineup, matching the output ceiling used across that lineup on the synchronous Messages API.
  • Adaptive Extended Thinking: Reasoning effort adjusts per request under an always-on adaptive mode, defaulting to high unless a caller sets it lower to save tokens.
  • Relaxed Cybersecurity Classifiers: Vetted Project Glasswing participants can use it to flag vulnerabilities directly in their own source code, a use case Anthropic's standard safeguards block on the general API.
  • Life Sciences Research Access: Biology and chemistry classifiers fire far less often on legitimate, benign requests than in the prior generation, per Anthropic, while dual-use questions still route to Anthropic's Opus models.
  • Multi-Platform Deployment: Runs under the model ID claude-mythos-5-1 on the Claude API and anthropic.claude-mythos-5-1 on Amazon Bedrock, plus Google Cloud Vertex AI and Microsoft Foundry.
  • Batch API Support: The asynchronous Message Batches API supports up to 300,000 output tokens with a beta header, at a discount versus the synchronous API.

Pros

  • Its strongest reasoning scores come from the independent ARC Prize foundation rather than Anthropic's own marketing claims.
  • Third-party leaderboards place it at 81.2% on SWE-bench Pro as of September 2026, an agentic coding benchmark Anthropic itself does not publish for this release.
  • Cybersecurity and biology safeguards are specifically tuned for legitimate defensive and research work rather than left at Fable 5.1's general-audience defaults.
  • Shares its full 1 million token context window and pricing with the generally available Claude Fable 5.1, so nothing is held back once access is granted.

Cons

  • Only reachable after a formal vetting process, which rules it out for developers who just want to try a Claude model on a whim.
  • Even approved users cannot use it for offensive work: it will not write a working exploit, run a penetration test, or scan a compiled binary for bugs.
  • At maximum reasoning effort it uses roughly 1.7 times the output tokens of the previous Fable 5 generation for the same task, raising cost per run.
  • Program enrollment accepts only US-headquartered applicants for now, so international teams cannot yet apply.

Benchmarks

  • ARC-AGI 2: 90% independent · 26 Sep 2026 — Abstract visual puzzles built to resist memorisation, % solved.
  • GDPval-AA v2: 1,853 vendor-reported · 26 Sep 2026 — Real knowledge-work deliverables judged against professionals, run by Artificial Analysis.
  • SWE-bench Pro: 81.2% independent · 26 Sep 2026 — Harder, longer real-repository coding tasks, % solved.
  • Humanity's Last Exam: 65% vendor-reported · 26 Sep 2026 — Expert-written questions across many fields, % correct.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

What are Claude Mythos 5.1's pricing plans in 2026?

Per-token rates match Claude Fable 5.1 exactly: $10 per million tokens in, $50 per million out. Writing a fresh prompt cache entry adds a premium ($12.50/M for a 5-minute entry, $20/M for a 1-hour one), while reading a cached entry drops to just $0.25/M, a 75% cut from the prior generation. Running requests through the async Batch API instead cuts both the input and output rate in half, and no free tier exists.

Who can access Claude Mythos 5.1 and how do you request it?

It is gated behind Project Glasswing, Anthropic's pair of vetted-access programs: the Cyber Verification Program for defensive security teams and the Life Sciences Verification Program for advanced biology research. An interested organization asks its Anthropic, AWS, or Google Cloud account representative to start enrollment, and for now only US-headquartered applicants are accepted while Anthropic works toward opening the program internationally.

What are the best alternatives to Claude Mythos 5.1?

Claude Fable 5.1 is the generally available version of the same underlying model, with Anthropic's standard cybersecurity and biology safeguards active and no program enrollment required. For general coding and reasoning work without any vetting process, Claude Sonnet 5 and Claude Opus 5.5 cover most use cases at a much lower per-token price and are available immediately through the standard Claude API.

Claude Mythos 5.1 or Claude Fable 5.1: which should you pick?

They are the same underlying model with an identical context window, output limit and pricing, so the choice comes down to safeguards and access. Pick Claude Fable 5.1 for the generally available version with Anthropic's standard safety classifiers active and no application process; pick Claude Mythos 5.1 only once your organization is a confirmed Project Glasswing participant and specifically needs those classifiers relaxed for legitimate defensive or research work.

What does it take to start using Claude Mythos 5.1?

Get your organization accepted into Project Glasswing first; that approval, not a signup form, is the real barrier to entry. From there it plugs into the same Messages API used across Claude 5, just under the model ID claude-mythos-5-1, so a codebase already calling Claude only needs a model-ID swap.

More AI Models on HokAI

Visit Claude Mythos 5.1 Official Page