Last updated: 2026-09-04
Fireworks AI is an inference API for open-source language, vision, and image models, positioned as an independent alternative to hyperscaler LLM hosting. It serves 10,000+ enterprise customers and processes over 13 trillion tokens a day through its proprietary FireAttention engine.
About Fireworks AI
Fireworks AI is an enterprise inference platform for open-source LLMs, built in 2022 by seven engineers from Meta's PyTorch team. The company raised $327 million in total funding, reaching a $4 billion valuation with its October 2025 Series C co-led by Lightspeed Venture Partners and Index Ventures. With $800 million in annualized revenue as of May 2026 and 10,000+ enterprise customers, Fireworks has established itself as the leading independent inference infrastructure company outside the major cloud providers.
The platform's speed comes from FireAttention, a proprietary CUDA kernel stack built end-to-end for transformer inference. It handles disaggregated serving, semantic caching, and speculative decoding internally, which keeps tail latency close to the median instead of the order-of-magnitude spikes common on shared inference stacks. Fireworks processes over 13 trillion tokens a day across a fleet of NVIDIA and AMD GPUs.
Fireworks is built for engineering teams running production, low-latency inference on open models without operating their own GPU fleet. Named customers include Cursor, Perplexity, Notion, Sourcegraph, Uber, DoorDash, Shopify, and Upwork. The platform ships managed fine-tuning (SFT, DPO, and reinforcement fine-tuning) for teams that want to specialize open models on their own data, and its OpenAI-compatible API added a beta Responses API with MCP tool support in 2026.
Screenshots


Pricing
Usage-priced across three products; no seat subscription. Free: $1 in credits to start, no subscription required. SERVERLESS INFERENCE bills per token on three dimensions - input, cached input (served from prompt cache, priced lower) and output - across Standard, Priority and Fast serving paths.
20. 60. Priority runs above Standard on the same models.
Batch inference is billed at 50% of serverless pricing on both input and output. 10 for Qwen3 8B. US-only models carry a 50% premium effective 1 September 2026.
1B-80B $3 / $6 / $6 / $12; 80B-300B $6 / $12 / $12 / $24; over 300B $10 / $20 / $20 / $40. Reinforcement fine-tuning is billed per GPU hour at on-demand rates. A Serverless Training API in private preview bills prefill, cached prefill, sample and train tokens separately.
Fine-tuned models serve at the same price as their base models. ON-DEMAND GPU is billed per GPU second with no charge for start-up time. 00.
5x premium. Enterprise deployments, with faster speeds, lower costs and higher rate limits, are quoted by sales, and volume, nonprofit, education and startup discounts are available.
| Tier | Monthly price | What it includes |
|---|---|---|
| Free credits | Free | $1 in free credits, no subscription and no card required to start. Self-serve signup. |
| Serverless Inference (pay per token) | Free | Per token on input, cached input and output, postpaid. Batch inference at 50% of serverless pricing. Embeddings billed on input only. US-only models carry a 50% premium from 1 Sep 2026. |
| Training (pay per training token) | Free | 00 (full-parameter DPO, models over 300B). Reinforcement fine-tuning bills per GPU hour at on-demand rates. Fine-tuned models serve at base-model prices. |
| On-Demand GPU (pay per GPU second) | Free | 00/hr. 5x premium. |
| Enterprise | Custom | Custom pricing via sales, for faster speeds, lower costs and higher rate limits. Volume, nonprofit, education and startup discounts available. |
Feature Comparison by Tier
| Feature | Serverless | On-demand GPU | Enterprise |
|---|---|---|---|
| Price | Pay per token, from $0.10 per 1M | Pay per GPU second, $8.00-$20.00/hr | Custom, quote only |
| Free credits to start | $1 | $1 | — |
| Self-serve signup | ✓ | ✓ | Contact sales |
| Setup and cold starts | Zero setup, no cold starts | No charge for start-up time | Managed |
| Serving paths | Standard, Priority, Fast | Dedicated GPU | Dedicated, higher rate limits |
| Batch inference discount | 50% of serverless pricing | n/a | Negotiated |
| Prompt-cache pricing | Cached input billed lower | n/a | Negotiated |
| Fine-tuning | Per 1M training tokens, from $0.50 | RFT billed per GPU hour | Negotiated |
| Fine-tuned model serving | Same price as base model | Same price as base model | Same price as base model |
| Region-restricted deployment | — | 1.5x premium | Negotiated |
| Rate limits | High, postpaid billing | Bounded by provisioned GPUs | Highest |
Key Features
- FireAttention Custom Inference Engine: Proprietary CUDA kernel stack delivers 167 t/s on DeepSeek V4 Pro, 5 times faster than DeepInfra and Novita at the same price, with P99 latency only 3.9x the median rather than typical 10x spikes.
- 400+ Open Model Catalog: Serves over 400 models including Llama 4, DeepSeek V4, Qwen 3, Mixtral, and FLUX, all accessible through an OpenAI-compatible API with no code changes needed for migration.
- Managed Fine-Tuning: Supports SFT, DPO, and reinforcement fine-tuning (RFT) with LoRA and full-parameter training; up to 100 LoRA adapters run simultaneously on a dedicated deployment at the same per-token rate as the base model.
- MCP Responses API (Beta): Fireworks Responses API connects open models to external tools via Model Context Protocol, handling the full agentic loop server-side with a single API call and SSE streaming.
- 99.8% Uptime at Scale: Independent monitoring shows 99.8% uptime in Q1 2026, the highest among specialized inference providers, sustaining roughly 180,000 requests per second across multi-cloud GPU orchestration.
- Zero Data Retention by Default: Fireworks does not log or store prompt or generation data for open models without explicit user opt-in, backed by SOC 2 Type II, HIPAA, GDPR, ISO 27001, and ISO 42001 certifications.
Pros
- Independently measured at roughly 5x faster than DeepInfra and Novita on DeepSeek V4 Pro at the same per-token price tier.
- Uptime among the highest of any independent inference provider in Q1 2026, per third-party monitoring, with audited SOC 2 Type II compliance.
- Fine-tuned deployments carry no extra per-token premium over the base model, unlike the typical 2-3x fine-tune markup competitors charge.
- OpenAI-compatible API means zero code changes for teams migrating from OpenAI or Together AI.
- Zero data retention by default for open models, with HIPAA BAA and ISO 27001 available for regulated industries.
Cons
- Image generation catalog is limited to roughly 5 models (FLUX and SDXL only) with no video generation support.
- Standard tier rate limits of 600 requests/minute require contacting sales to scale, unlike Together AI which exposes higher self-serve limits.
- Documentation is sparse for advanced configurations; no quickstart guide exists for fine-tuning workflows as of mid-2026.
- Batch inference discounts lag behind Together AI and OpenAI, whose batch pricing programs have been live longer.
Data Handling
- Training-data policy
- Does not log or store prompt or generation data for open models without explicit user opt-in. Sensitive training data never persists on the platform.
- Data retention
- Zero retention
- Compliance
- SOC 2 Type II · HIPAA · GDPR · ISO 27001 · ISO 27701 · ISO 42001 · FedRAMP · CSA Star Level 1
Frequently Asked Questions
What does Fireworks AI actually cost?
Fireworks AI bills purely pay-per-token, with no monthly subscription or seat fees. Serverless inference starts at $0.20 per million tokens for 8B-class models and $0.90 per million tokens for 70B-class models, with batch inference at 50% of those rates. Dedicated on-demand GPUs are $8.00 an hour for an H100 or H200 and $13.00 an hour for a B200 from 1 September 2026. Volume, nonprofit, education, and startup discounts are available on request.
Does Fireworks AI have a free plan?
Not a permanent one. Every new account opens with $1 in free credits, roughly a thousand calls to an 8B-class model, then switches to standard pay-per-token billing with no minimum commitment. Nonprofits get 40-80% off standard rates and educational institutions 50-90%, both through an application-based program.
What are Fireworks AI's closest competitors?
Groq, Together AI, and Replicate come closest. Groq's custom LPU chips beat Fireworks on raw single-model speed but skip fine-tuning and cover a much smaller catalog, Together AI carries a broader spread of open and MoE models, and Replicate aims at quick prototyping plus image or video generation rather than high-throughput production. Fireworks tends to win the teams that want fine-tuning, compliance, and enterprise throughput on one platform.
Fireworks AI or Groq: which should you pick?
Groq's custom LPU silicon can out-pace Fireworks on raw tokens-per-second for the handful of models it supports, but its catalog covers a fraction of what Fireworks serves and it has no fine-tuning offering at all. Fireworks answers with a much larger open-model catalog, full SFT/DPO/RFT fine-tuning with LoRA adapters, and SOC 2 Type II plus HIPAA compliance that Groq does not match. For teams that need raw single-model speed above everything else, Groq wins; for teams that need model variety, fine-tuning, and regulated-industry compliance, Fireworks is the stronger pick.
How do you set up Fireworks AI?
Register at fireworks.ai, which opens the account with $1 in credits, then issue an API key from the dashboard. Because the request and response format is a drop-in match for OpenAI's, you can install the OpenAI Python or Node SDK and point the base URL at the Fireworks endpoint. Choose a model ID from the catalog, one of the Llama or DeepSeek entries for instance, and send a chat completion; existing OpenAI code usually runs unchanged.
Top Alternatives
HokAI guides covering Fireworks AI
- Fireworks AI vs Together AI: Which Should You Use in 2026?: Fireworks and Together AI now charge the identical per-token rate on DeepSeek V4 Pro, so price no longer separates these two open-model inference platforms.
- Best Generative AI Infrastructure in 2026: 7 Platforms, Three Separate Decisions: Fireworks AI, Together AI, Groq, Modal, Replicate, Archil and Cast.ai compared on price and speed, plus the storage and cost layers most 2026 roundups skip.
- Groq's Llama Pricing Just Went Enterprise-Only. Where Does That Leave Fireworks AI?: Groq moved Llama to Enterprise-only pricing in Sept. 2026. Compare real speed, cost, and fine-tuning against Fireworks AI before choosing an inference API.
- What Together AI Is Actually For, Now IBM Is Betting $240 Million on It: IBM is spending $240M on Together AI's inference cluster and Together raised $800M in July. Here's what changed for buyers this week, and what did not.
- Hugging Face vs OpenPipe: Which Should You Use in 2026?: OpenPipe's SaaS was absorbed into CoreWeave in 2025 and Hugging Face retired AutoTrain. Here's what's actually left to compare for fine-tuning in 2026.