All AI guides
Comparison8 min read

Groq's Llama Pricing Just Went Enterprise-Only. Where Does That Leave Fireworks AI?

Groq is an inference API built on custom LPU chips, priced per token but now selling only OpenAI's GPT-OSS and Whisper models publicly. Fireworks AI is a broader inference and fine-tuning platform hosting over 400 open models, including DeepSeek, Kimi and GLM, at published per-token rates.

The short version

Groq still wins on raw speed for the models it prices publicly, but it just moved Llama 3.1 8B and 3.3 70B to Enterprise-only pricing. Fireworks AI wins once you need more than a few models, fine-tuning, or DeepSeek and Kimi-class context windows that Groq does not host at all.

Every third-party comparison of Groq and Fireworks AI published this year quotes the same number for Llama 3.3 70B on Groq: 59 cents per million input tokens, 79 cents per million output tokens. Groq's own developer docs, checked on 3 September 2026, list that same model under a different heading entirely: Enterprise, contact sales.

That one line change reframes the whole decision. Groq and Fireworks AI both sell access to open-weight language models over an API, and both promise a speed a general-purpose GPU cloud cannot match. But Groq just narrowed what you can actually buy at a fixed price, while Fireworks kept expanding its catalog. For a team choosing between them in September 2026, the axis that matters most is no longer only tokens per second. It is which models each vendor will still sell you off the shelf.

The verdict

Pick Groq if the model your product runs on is still on its public price list, OpenAI's GPT-OSS or Whisper, and raw generation speed is the product itself: a voice agent, a live coding assistant, anything where a user notices 200 milliseconds. Pick Fireworks AI if you need more than a handful of models, want DeepSeek V4 Pro-class context windows, or plan to fine-tune anything at all. Groq does not offer fine-tuning, on any model, at any price.

What changed this month

Groq's pricing has always been narrow by design: a short list of models running on custom silicon, priced per token, no negotiation required. As of this check, that list got shorter. Llama 3.1 8B and Llama 3.3 70B, the two models nearly every 2026 comparison of Groq cites for its headline cheap-and-fast pricing, now carry "Enterprise (contact sales)" instead of a rate card, according to Groq's own developer documentation.

OpenAI's GPT-OSS 120B and GPT-OSS 20B kept their public rates: $0.15 input and $0.60 output per million tokens for the larger model, $0.075 and $0.30 for the smaller one.

Groq's Supported Models page showing Llama 3.1 8B and Llama 3.3 70B both marked Enterprise with a price column reading Contact Sales, next to GPT-OSS 120B's public speed and capability card

Groq's own developer docs, captured 3 September 2026. Llama's price column now reads Contact Sales instead of a per-token rate.

The practical effect is simple. A developer who signs up today expecting to self-serve Llama on Groq, the way most guides published in the last eight months describe it, hits a sales form instead of a checkout page. Whether this is a temporary capacity move or a permanent repricing is not stated anywhere in Groq's documentation, and Groq has not published a changelog entry explaining the change.

Price, model by model

Where Groq still publishes a rate, it undercuts most GPU-based competitors on paper. GPT-OSS 20B costs $0.075 for a million tokens in, $0.30 for a million tokens out. Whisper Large V3 Turbo transcribes audio at $0.04 per hour, a fraction of the standard Whisper Large V3 rate of $0.111 per hour on the same platform.

Fireworks prices per model rather than per tier. DeepSeek V4 Pro costs $1.32 per million input tokens and $3.96 per million output tokens on Fireworks' own pricing page, with a faster, cheaper DeepSeek V4 Flash variant at $0.22 and $0.66. New signups get $1 in free credits, enough for a short test but not a real load test.

Neither company states what its enterprise floor actually costs. That is the gap this article cannot close: Groq will not say what "contact sales" means in dollars for Llama, and Fireworks' dedicated-GPU tier is quoted only as a $7 to $20 per hour range depending on GPU type, with the exact figure locked behind a sales call too.

Fireworks AI's pricing page headline reading Pricing to scale from idea to enterprise, with a self-serve Get Started button next to a separate Contact Us path

Fireworks' pricing page, captured 3 September 2026. Self-serve stays the default; enterprise terms sit behind a sales conversation, the same split Groq now uses for Llama.

Free-tier access also differs in shape, not just size. Groq's free plan needs no credit card, but caps usage per model: roughly 30 requests per minute and up to 14,400 requests per day on its higher-volume models, according to Groq's own rate-limit documentation, with a paid Developer tier raising those ceilings roughly tenfold and cutting on-demand pricing by 25 percent.

Fireworks' $1 in free credits works differently: it is a spending cap, not a request-rate cap, so a handful of calls against a large model burns through it far faster than the same handful of calls against GPT-OSS 20B on Groq.

Speed: what an independent benchmark actually shows

Vendor-published tokens-per-second numbers deserve some caution, since both companies measure on their own hardware under their own load conditions. Artificial Analysis, a third party that benchmarks inference providers independently, measured Groq's fastest tracked model, GPT-OSS 20B running in low-latency mode, at 927.5 tokens per second. It measured Fireworks' DeepSeek V4 Pro at 97 tokens per second, and the faster DeepSeek V4 Flash variant at 261 tokens per second on the same platform.

That is not an apples-to-apples number. Groq's figure is a small, dense model on purpose-built LPU silicon; Fireworks' figure is a much larger model on general-purpose GPUs. But it shows the honest shape of the tradeoff. The model Fireworks runs fastest is still roughly a quarter the speed of Groq's fastest model, and Groq cannot run DeepSeek V4 Pro at any speed, because it is not in Groq's catalog at all.

Model breadth and fine-tuning

This is where Fireworks' pitch stops being about speed and starts being about not having to choose. Its catalog runs past 400 models, including several with no Groq equivalent: Moonshot's Kimi K3, a million-token-context model priced at $3 per million input tokens and $15 per million output tokens, and Zhipu AI's GLM 5.3, also at a million tokens of context, priced at $1.40 and $4.40.

Fireworks also sells fine-tuning directly, with published per-token training rates instead of a quote. LoRA fine-tuning on a model up to 16 billion parameters costs $0.50 per million training tokens for supervised fine-tuning and $1.00 for DPO; full-parameter fine-tuning at the same size tier runs $1.00 and $2.00.

The rate climbs with model size, up to $20 to $40 per million tokens for full-parameter fine-tuning on models over 300 billion parameters. Groq has no equivalent product at any price. There is no fine-tuning API, managed or otherwise, anywhere in Groq's documentation.

Where Groq wins

If a product's bottleneck is genuinely generation latency, and the model fits Groq's shrinking public catalog, nothing on Fireworks, Together AI, or any GPU-based platform matches it. A voice agent generating a spoken response, or an IDE autocomplete feature where every extra 100 milliseconds is felt directly, is the specific use case Groq was built for. It still wins that case outright.

Where Fireworks wins

The moment a team needs a second or third model, needs to fine-tune anything, or wants a model Groq does not host at all, the decision stops being close. Fireworks AI's fine-tuning pricing, its DeepSeek and Kimi catalog, and its willingness to sell all of it at a published per-token rate make it the default for any team that has not locked in on one specific open model yet. Modal covers a similar training-first niche with more infrastructure control, but does not compete on Fireworks' breadth of pre-hosted models.

Fireworks' bet is that most teams do not know their final model choice on day one. A support-ticket classifier might start on a small, cheap model and graduate to a fine-tuned version of something larger once real traffic shows where the small model fails. That whole path, from stock model to fine-tuned production deployment, stays inside one Fireworks account and one bill. Rebuilding it on Groq is not an option, since fine-tuning is not a Groq product at any tier.

The turn

The obvious rebuttal here: raw tokens per second rarely bottlenecks a real product, because network latency, application logic and rate limits dominate what a user actually experiences long before inference speed does. That holds for a chat interface, where a half-second delay is invisible inside a longer response.

It stops holding for two specific workloads: real-time voice agents, where every added millisecond of generation time is audible and not just measurable, and high-volume batch pipelines, where a roughly ninefold throughput gap between GPT-OSS on Groq and DeepSeek V4 Pro on Fireworks becomes a roughly ninefold difference in compute spend for the same job.

Switching cost

Both platforms use OpenAI-compatible chat completion APIs, so a working integration usually ports in an afternoon: change the base URL and the model name, and most client libraries need nothing else. What does not port is anything built around Groq's speed budget.

A voice pipeline tuned around Groq's sub-100-millisecond generation time needs its buffering and timeout logic rewritten if it moves to a GPU-based platform running four to nine times slower, no matter how similar the API looks on paper. Teams weighing either option against a wider field, including model choice, seat cost and use case fit, can run the comparison through Smart Match rather than reading every vendor's pricing page by hand.

What to watch

Groq's shift to Enterprise-only pricing on its two most requested Llama models is new enough that it may not hold. If Groq restores a public rate for Llama, or expands its catalog past its current narrow list, the speed-versus-breadth tradeoff described here narrows again. Watch Groq's developer documentation, not its marketing pricing page, for the next change. This run found Groq's public marketing page did not render pricing details at all; only the developer docs did.

Frequently asked questions

Is Groq faster than Fireworks AI?

For the models each platform still prices publicly, yes. Artificial Analysis measured Groq's fastest tracked model at 927.5 tokens per second, against 97 tokens per second for Fireworks' DeepSeek V4 Pro. The two numbers are not the same model on both platforms, since Groq does not host DeepSeek V4 Pro at all.

Can I fine-tune a model on Groq?

No. Groq has no fine-tuning API, managed or otherwise, on any model. Fireworks AI sells LoRA and full-parameter fine-tuning directly, priced per training token and scaled by model size.

Why does Groq list Llama pricing as Enterprise, contact sales?

Groq's own developer documentation moved Llama 3.1 8B and Llama 3.3 70B off public per-token pricing as of this check on 3 September 2026. Groq has not published a changelog entry explaining whether the change is permanent.

What models does Fireworks AI host that Groq does not?

Fireworks' catalog runs past 400 models, including DeepSeek V4 Pro, Moonshot's Kimi K3 and Zhipu AI's GLM 5.3, none of which appear in Groq's smaller, speed-focused lineup. Groq's public list stays limited to a handful of models, chosen for raw throughput rather than catalog breadth.

Do Groq and Fireworks AI use the same API format?

Both expose OpenAI-compatible chat completion endpoints, so switching usually means changing the base URL and model name. Anything tuned around Groq's low-latency generation speed still needs its timeout and buffering logic rewritten for a GPU-based platform.

Covered in this guide

  • Groq: Fast, low cost inference powered by Language Processing Units (LPUs)
  • Fireworks AI: Enterprise LLM inference platform from Meta PyTorch veterans: 400+ open models, 167 t/s on DeepSeek V4 Pro, pay-per-token pricing, 99.8% uptime.
  • DeepSeek V4 Pro: DeepSeek V4 Pro: 1.6T-param open-source MoE (April 2026), 80.6% SWE-bench Verified with 1M token context under MIT license. $1.74/$3.48 per 1M tokens.
  • GLM 5.3: GLM-5.3 arrived in August 2026 as Z.ai's coding- and cybersecurity-focused update, post-trained on the same base as its predecessor.
  • Kimi K3: 2.8T-parameter open-weight MoE model from Moonshot AI (July 2026) with a 1M-token context window and 93.5% GPQA Diamond, the top open score.
  • Modal: AI infrastructure that developers love: serverless compute for ML inference, training, and batch processing

Sources

Still deciding?

This guide covers a handful of options. Smart Match checks every listing in the directory against how you actually work and what you can spend, then hands you the shortlist and the reason behind each pick.

Start Smart Match

Related guides

All AI guidesBrowse the AI directory