Muse Spark 1.1: 1M Context & Pricing (2026)
Meta's first paid AI model (July 2026): a 1M-token context window, native computer use, zero-shot tool and MCP support, and pay-as-you-go API pricing per token.
Released three months after the original Muse Spark, it trails top execution-focused coding models by roughly 16 points on a leading terminal-execution benchmark, but wins on cost-per-completed-task for teams running multi-app computer-use agents rather than chasing the single highest coding score available today.
Meta Superintelligence Labs' multimodal reasoning model for agentic and computer-use tasks scores 62.1 on Artificial Analysis' tool-enabled Intelligence Index, its main differentiator from the original Muse Spark. It runs a self-managing 1M-token context window and is Meta's first monetized model API, sold through the Meta Model API.
Provider: Meta AI · Family: Muse
Context window: 1,048,576 tokens · Max output: 131,072
Input modalities: text, image, video, audio, pdf, tool-calls · Output: text, tool-calls
About Muse Spark 1.1
Muse Spark 1.1 is the second model in Meta Superintelligence Labs' Muse family, released on July 9, 2026, three months after the original Muse Spark debuted on April 8, 2026 as Meta's first closed-weight frontier model. Where Muse Spark 1.0 was a general-purpose reasoning model competing on Chatbot Arena, Muse Spark 1.1 is purpose-built for agentic work: tool use, computer control, and multi-agent orchestration. It marks Meta Superintelligence Labs' most significant strategic pivot yet under Chief AI Officer Alexandr Wang, moving Meta from a free, open-weight default to its first paid developer API. On the Artificial Analysis Intelligence Index with tools enabled, Muse Spark 1.1 scores 62.1, ahead of Claude Opus 4.8 (57.9), GPT-5.5 (52.2), and Gemini 3.1 Pro (51.4), and a clear jump from Muse Spark 1.0's 50.4 on the same tool-augmented measure. On execution-focused coding benchmarks it trails the frontier: a Terminal-Bench 2.0 score of 59.0 sits 16 points behind GPT-5.4's 75.1, and SWE-bench Pro (a harder, contamination-resistant successor to SWE-bench Verified) puts it at 61.5%. It leads on HealthBench Hard at 42.8, ahead of GPT-5.4 (40.1), Gemini 3.1 Pro (20.6), and Grok 4.2 (20.3). Meta has not published exact GPQA Diamond or MMLU-Pro percentages for this release; third-party trackers rank it around 12th and 9th respectively among frontier models tracked in mid-2026, but no verified raw score exists for either as of this writing. The context window runs to 1,048,576 tokens, self-managed with active compaction so long agent sessions degrade gracefully instead of hitting a hard wall. Independent trackers disagree on the maximum output limit: EmpirioLabs reports 131,072 tokens per response while Vals AI's index lists 256,000, a gap that likely reflects different test-harness configurations rather than a single documented API ceiling. Meta reports a time-to-first-token of 1.31 seconds and throughput around 120 tokens per second for the xhigh reasoning variant, both competitive against similarly priced reasoning models. The model is natively multimodal, accepting text, images, video, audio, and PDF documents and returning text and tool calls. Its headline capability is tool orchestration: Meta trained it to generalize zero-shot to new native tools, MCP servers, and custom skills, to run as the primary agent that plans and delegates to parallel subagents, or as a subagent that knows when to escalate back to a main agent. Its computer-use mode reads and acts across multiple applications inside one session, weighing scripting, direct UI clicks, and batched actions at each step; Meta's own demo has it pulling product photos out of a smartphone video and posting a Facebook Marketplace listing entirely unattended. Meta prices the new Model API well under what Anthropic and OpenAI charge for their flagship reasoning models, positioning it for high-volume agentic and coding workloads rather than occasional chat use. New developer accounts also start with a batch of free usage credit, and the underlying model remains free to try for consumers through Thinking mode in the Meta AI app and at meta.ai, no API key required. Exact per-token rates and the free-credit amount are broken out in the pricing FAQ below. The Meta Model API launched in public preview on July 9, 2026, restricted to US developers. As of this writing there is no confirmed listing on AWS Bedrock, Google Vertex AI, or Azure, and third-party aggregator access has been rolling out unevenly to eligible US accounts rather than shipping as a full marketplace listing on day one. Meta has not disclosed this model's parameter count, architecture (dense versus mixture-of-experts), or training data cutoff. Rather than publishing a traditional system card, Meta released a Muse Spark 1.1 Evaluation Report documenting safety testing under its Advanced AI Scaling Framework across three catastrophic-risk domains: Chemical and Biological, Cybersecurity, and Loss of Control. Meta reports a multi-layered mitigation strategy that rejects prompts seeking Chemical or Biological weapons uplift, re-validated specifically for the API deployment surface, and states the model operates within safe margins across all three domains. Muse Spark 1.1 is best understood as a cost-efficient orchestration model rather than a single benchmark-topping generalist. Teams building multi-app computer-use agents, parallel subagent pipelines, or high-volume agentic coding workloads get a markedly cheaper alternative to Claude Opus 4.8 or GPT-5.5 with competitive tool-use scores. Teams that need the single highest coding-execution benchmark, a full vendor system card for compliance review, or API access outside the US today are better served by GPT-5.4, Claude Opus 4.8, or Gemini 3.1 Pro until Meta widens availability.
Pricing
$1.25 per 1M input tokens and $4.25 per 1M output tokens on the Meta Model API, launched as a US-only public preview. New accounts start with $20 in free usage credit before pay-as-you-go billing kicks in.
Key Features
- 1M-Token Self-Managing Context: A large working memory that quietly compacts itself as an agent session runs long, instead of erroring out at a hard token limit.
- Native Computer Use: Reads a screen and takes real actions across several apps in one run, switching between scripts, clicks, and batched steps as needed.
- Zero-Shot Tool and MCP Support: Picks up brand-new tool definitions, MCP servers, and custom skills at inference time without any task-specific fine-tuning.
- Parallel Subagent Orchestration: Can run the show as the lead planner splitting work across subagents, or take orders as a subagent that escalates when it's stuck.
- Meta Model API: Meta's first monetized developer API, opened in US public preview with straightforward pay-as-you-go token billing and a starter credit.
Pros
- Prices well under flagship rivals on a per-token basis while staying competitive on tool-augmented intelligence benchmarks.
- Tops every frontier model Artificial Analysis benchmarked on HealthBench Hard as of July 2026.
- Context window manages itself across long agent sessions instead of failing at a hard cutoff.
Cons
- Trails top execution-focused coding models by a wide margin on terminal-style coding benchmarks.
- Ships a Preparedness/Evaluation Report instead of a traditional per-model system card.
- API access is a US-only public preview with no confirmed AWS Bedrock, Vertex, or Azure listing.
Benchmarks
- swe bench pro: 61.5
- healthbench hard: 42.8
- terminal bench 2: 59
- artificial analysis intelligence index: 62.1
- artificial analysis speed tokens per sec: 120
Frequently Asked Questions
How much does Muse Spark 1.1 cost per 1M tokens?
Meta's Model API charges $1.25 per 1 million input tokens and $4.25 per 1 million output tokens, well under the $5 input and $25 to $30 output rates Anthropic and OpenAI charge for their comparable flagship models. New API accounts also start with $20 in free usage credit. Consumer access through the Meta AI app's Thinking mode stays free with no API key required.
How does Muse Spark 1.1 compare on benchmarks vs GPT-5.5?
It leads GPT-5.5 on the tool-augmented Artificial Analysis Intelligence Index, reflecting Meta's orchestration-focused training for this release. On raw execution coding, a related OpenAI model beats its terminal-execution benchmark score by a wide margin, so which model wins depends heavily on the task.
Is Muse Spark 1.1 open source or proprietary?
It is proprietary and closed-weight, a departure from Meta's open-weight Llama family. Access is limited to the paid Meta Model API (public preview, US developers) or free consumer chat in the Meta AI app; there is no downloadable checkpoint or open license.
Does Muse Spark 1.1 train on user data?
Meta had not published specific data retention or training opt-out terms for the Meta Model API as of its public preview launch. Teams needing a documented zero-retention policy should confirm current terms directly with Meta before sending production traffic.
Who is Muse Spark 1.1 best for and who should avoid it?
It suits teams building cost-sensitive, multi-app computer-use agents and parallel subagent pipelines, given its self-managing context window and pricing well below flagship rivals. Teams chasing the single highest coding-execution score, or needing a full vendor system card for compliance, should look elsewhere in the frontier lineup instead.
Top Alternatives
- GPT-5.5: Pick this model if you want much lower per-token pricing for agentic and tool-use workloads.
- Claude Opus 4.6: Pick this model over the rival Opus release when cost-per-completed-task on multi-app automation matters more than peak reasoning benchmarks.
- Gemini 3.1 Pro: Pick the rival Pro release instead if you need a stronger execution-focused coding score today.