What Together AI Is Actually For, Now IBM Is Betting $240 Million on It
Together AI is a hosted inference platform for open-weight AI models, sold as serverless per-token API calls, reserved Provisioned Throughput, or dedicated GPU rental. As of August 2026 it charges $1.04 per million tokens for Llama 3.3 70B and $0.14 to $0.28 for DeepSeek V4 Flash, unchanged by its recent funding and IBM infrastructure deal.
The short version
Together AI runs inference for open models like Llama, DeepSeek and Kimi K3. IBM just committed $240 million to a dedicated cluster live in Q1 2027, weeks after Together's $800 million funding round. Neither deal changed today's per-token prices, where Fireworks still undercuts Together on Llama 3.3 70B.
IBM will spend $240 million putting Nvidia's newest chips under Together AI's control, with the cluster live by the first quarter of 2027.
That single line, from an August 11 joint announcement, is the second nine-figure signal Together AI has sent this summer. On July 1 the company closed an $800 million funding round at an $8.3 billion valuation. Neither event changes what a five-person startup pays per token today. What it changes is who Together AI can now afford to serve, and on what hardware, starting next year rather than this one.
What Actually Changed
The IBM deal is a multi-year agreement, confirmed by IBM's newsroom, to deploy Nvidia HGX B300 systems with Spectrum-X networking on IBM's cloud. Both companies call it the first dedicated large-scale inference cluster of that kind on the platform. Together AI operates it, IBM hosts it, and Nvidia supplies the chips.
"Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale," the company's CEO Vipul Ved Prakash said in the announcement. Availability is targeted for Q1 2027, not now.
The funding round landed six weeks earlier. Aramco Ventures led the $800 million raise, with Nvidia among the other backers. The company said bookings crossed $1.15 billion in the prior quarter and that new investors pledged over 500 megawatts of future compute capacity. It named Decagon as one paying customer, which cut its inference bill sixfold after migrating onto the platform. That is a real number, but it describes a past migration, not this week's price sheet.
The same stretch brought two smaller, more immediate moves. It became a launch partner for Moonshot AI's Kimi K3, a 2.8 trillion parameter open model with a 1 million token context window, serving it on day zero in late July. It also stood up a dedicated GPU cluster for YC's portfolio companies, letting those startups rent short-term capacity instead of signing long-term contracts.
The Three Ways You Actually Buy This
Strip away the funding news and the product itself is unchanged: three tiers, priced per unit of a different thing. Serverless inference bills per token, with no GPU provisioning, the default for prototyping. Provisioned Throughput reserves capacity at a flat $0.05 per unit per minute, meant for production traffic that needs a guaranteed floor. Dedicated inference rents a GPU outright, at $5.49 an hour on demand for an H100, for teams with steady enough load to make that math work.
The token prices vary a lot by model. On Together's own pricing page, Llama 3.3 70B runs $1.04 per million tokens for input and output alike, and Kimi K3 runs $3.00 in, $15.00 out. DeepSeek V4 Flash is the cheap end, $0.14 in and $0.28 out, with cached input dropping to $0.03. None of those four numbers moved after the round closed.
Who The New Money Helps
The IBM deal is aimed at enterprise buyers who want an inference vendor with a name their procurement team already trusts on the invoice, not at a startup choosing between APIs this afternoon. It gives IBM's cloud customers a path to open-weight models without running their own GPU fleet, and it gives the company a distribution channel it did not have in July.
Model makers benefit too. Moonshot AI got a launch partner with real serving capacity behind Kimi K3 on day zero, the kind of deal that used to require its own infrastructure build. YC's portfolio gets GPU access without the multi-year commitment that smaller teams usually cannot negotiate.
Who It Leaves Behind
Nobody in that list is a four-person team picking between Fireworks AI and this platform on price alone, and that comparison has not moved. Fireworks's own pricing documentation lists a flat $0.90 per million tokens for models above 16 billion parameters, the bracket Llama 3.3 70B falls into, against $1.04 here. On DeepSeek V4 Flash the two hosts are identical, $0.14 in and $0.28 out on both, so the gap is model-specific rather than a blanket discount, and neither company's funding shows up in either number.
Replicate's marketplace model, where independent providers compete on the same model rather than one company owning the stack, is untouched by any of this. Teams that picked it for image, video or niche model coverage rather than raw LLM throughput have no reason to reconsider based on an enterprise cloud deal aimed at chat and code models. Self-hosted vLLM setups are further removed still: capacity news at a hosted vendor has zero bearing on a GPU cluster you already own.
The Case Against Waiting
The obvious counterargument is that more capital, and specifically Nvidia writing a check, signals which vendor survives the next downturn, so locking in now is the safer bet even at today's slightly higher price. That case has some force for a team planning three years out. It has none for a team choosing a host this quarter.
The B300 cluster is not live until Q1 2027. Nothing in either announcement lowered a single price on the company's own page, and Fireworks still beats it on the exact model most teams start with. A funding round is evidence a company can keep operating. It is not evidence about this week's bill.
What To Watch
The number worth tracking is not the valuation. It is whether Together AI moves its serverless price sheet before the IBM cluster ships. A cut to the Llama 3.3 70B rate, or a public SLA upgrade on Provisioned Throughput, would be the first sign the new capital is reaching this week's customers rather than next year's. Watch Nvidia's B300 shipment schedule too. A cluster live ahead of Q1 2027 would be the clearest evidence the wait-and-see case is wrong.
Until one of those happens, the $240 million and the $8.3 billion valuation belong to a different decision than the one in front of a team picking an API today.
Frequently asked questions
What is Together AI used for?
Together AI hosts open-weight models such as Llama, DeepSeek and Kimi K3 behind a single API, so developers can call them without provisioning GPUs. It offers three buying modes: per-token serverless calls, reserved Provisioned Throughput, and dedicated GPU rental for steady production traffic.
Is Together AI cheaper than Fireworks AI?
It depends on the model. Fireworks charges a flat $0.90 per million tokens for models over 16 billion parameters, versus $1.04 on Together AI for Llama 3.3 70B. On DeepSeek V4 Flash the two platforms charge the identical $0.14 input and $0.28 output rate, so there is no blanket discount either way.
What does the IBM deal actually change for a Together AI customer?
Very little right now. The multi-year, $240 million agreement deploys Nvidia HGX B300 systems on IBM Cloud, but availability is targeted for Q1 2027. It gives IBM Cloud customers a route to open-weight inference and gives Together AI enterprise distribution, without altering today's serverless pricing.
How much did Together AI raise in its Series C?
Together AI raised $800 million on July 1, 2026, led by Aramco Ventures with participation from Nvidia, Vista Equity Partners and General Catalyst, at an $8.3 billion valuation. The company said quarterly bookings had crossed $1.15 billion and new investors pledged over 500 megawatts of future compute.
What is Provisioned Throughput on Together AI?
Provisioned Throughput is Together AI's reserved-capacity tier, billed at a flat $0.05 per unit per minute rather than per token. It is built for production workloads that need a guaranteed throughput floor, sitting between pay-as-you-go serverless calls and renting a dedicated GPU outright.
Covered in this guide
- Together AI: The AI Native Cloud—full-stack platform for training, fine-tuning, and deploying open-source AI models
- Fireworks AI: Enterprise LLM inference platform from Meta PyTorch veterans: 400+ open models, 167 t/s on DeepSeek V4 Pro, $0.20/1M tokens, 99.8% uptime.
- Replicate: Run open-source AI models via API without managing any GPU infrastructure.
- Kimi K3: Native multimodal AI assistant with Agent Swarm capabilities and 256K context window
- Moonshot AI: Beijing AI lab founded in 2023 by Yang Zhilin, maker of Kimi and the open-weight Kimi K2 models, valued at $20B in 2026.
- Nvidia: Founded 1993, NVIDIA is the world's most valuable company (~$4.85T, July 2026), building the GPUs, CUDA stack, and open Nemotron models that run most of the AI industry.
Sources
- Together AI Pricing
- IBM and Together AI Sign Multi-Year Agreement
- Announcing our $800M Series C to accelerate the shift to open-source AI
- Serverless Pricing
- Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community
- Together AI announces strategic partnership with Moonshot AI
Still deciding?
This guide covers a handful of options. Smart Match checks every listing in the directory against how you actually work and what you can spend, then hands you the shortlist and the reason behind each pick.
Start Smart MatchRelated guides
- The AI Tool Ecosystem in 2026: Buy the Meter, Not the Category
- CometAPI vs Orthogonal: Same Gateway Label, Two Different Jobs
- xAI Spent $60 Billion on Cursor. Grok Build Now Runs Two Coding Models.
- OpenAI Killed Sora in March. Five Months Later, Here's Who Actually Won.
- Enterprises Still Can't Govern Their AI Agents. Brussels Just Moved the Deadline.
- What Qwen Is Actually For, Now It's Priced Below Claude and GPT-5.6