Together AI

The AI Native Cloud—full-stack platform for training, fine-tuning, and deploying open-source AI models

Together AI · Free tier available

Last updated: 2026-07-08

Together AI is an AI-native cloud platform serving over 100 billion tokens daily across 200+ open-source models including Llama 4, DeepSeek-R1, Qwen 3, and Mistral. The platform covers the full ML stack: serverless inference, dedicated GPU endpoints, GPU cluster rental, fine-tuning, audio TTS/STT, image generation, and video generation. It holds SOC 2 Type II, HIPAA, and ISO 27001:2022 certifications, and raised an $800M Series C in July 2026 at an $8.3B valuation.

About Together AI

Together AI is a research-driven cloud platform that empowers developers and enterprises to build, train, fine-tune, and deploy open-source AI models at scale. Founded in 2022, the company provides a comprehensive AI infrastructure layer offering serverless inference, dedicated model deployment, batch processing, fine-tuning, GPU clusters, and managed storage—all optimized with cutting-edge research breakthroughs like FlashAttention, Medusa, and speculative decoding. The platform supports 200+ open-source models including Llama, Mistral, Qwen, DeepSeek, and proprietary models like Mamba-3 and Dragonfly. Together AI's cost-effective token-based pricing and research-optimized infrastructure deliver 2x faster inference and 60% lower costs compared to alternatives. The company is backed by $534M in funding (Series B in Feb 2025 at $3.3B valuation) from investors including NVIDIA, Salesforce Ventures, Kleiner Perkins, and General Catalyst. Unlike closed-source alternatives, Together AI prioritizes open-source transparency and data privacy. Teams can deploy models on serverless infrastructure for variable workloads, dedicated endpoints for production scale, or GPU clusters for custom training—all without vendor lock-in. The platform powers AI startups like Cursor, Pika Labs, and NexusFlow in production.

Pricing

Serverless inference from $0.03/1M input tokens (smallest models) to $4.50/1M output tokens (largest frontier models). Batch API at up to 50% discount. Fine-tuning from $0.48/1M training tokens (models up to 16B; $4 minimum per job), $1.50-$4.12/1M (17B-69B), $2.90-$8.00/1M (70B-100B). GPU clusters on-demand: H100 $3.99/GPU-hour, H200 $5.99/GPU-hour, B200 $8.19/GPU-hour; reserved pricing from $3.09/hr (H100 181+ days). Dedicated inference endpoints: H100 HGX $5.49/hr, B200 $8.99/hr. Audio transcription from $0.0015/min. Image generation from $0.0006/image. Video generation from $0.14/video. Free tier with starter API credits for new accounts.

Key Features

  • Serverless Inference: Deploy 200+ open-source models on demand with pay-per-token pricing, no infrastructure management, and no long-term commitments. Models include Llama 4, DeepSeek R1 and V3, Qwen 3, and Mistral.
  • Dedicated Model Inference: Deploy models on single-tenant reserved GPU compute with guaranteed throughput, full isolation, and production SLA. Available on H100 HGX and B200 hardware for latency-critical workloads.
  • Fine-Tuning at Scale: Fine-tune open-source models from 8B to 405B parameters using SFT, DPO, and LoRA, with 6x higher throughput than alternatives per company benchmarks. Users retain full ownership of fine-tuned weights.
  • GPU Clusters and Accelerated Compute: Rent H100, H200, or B200 GPU clusters on-demand from 16 GPUs to thousands, with external OIDC authentication for team access and automatic node repair for production reliability.
  • Batch Processing API: Process large asynchronous inference workloads up to 30 billion tokens per model at up to 50% lower cost than real-time serverless, with flexible scheduling and error recovery.
  • Audio AI (TTS and STT): Real-time text-to-speech via WebSocket with Orpheus 3B and Kokoro 82M models. Real-time speech-to-text transcription using Whisper via WebSocket streaming. Both TTS and STT support REST and streaming endpoints.
  • Image and Video Generation: Serverless image generation via FLUX.1 and Google Imagen 4.0 Ultra. Video generation via Seedance 2.0 (text-to-video, image-to-video, 4K up to 3840x2160) and Google Veo 3.0 with audio.
  • Managed Storage and Code Sandboxes: High-performance shared filesystem storage colocated with compute at $0.16/GiB/month with zero egress fees. Secure code sandboxes for AI agent development at $0.0446/vCPU-hour.

Pros

  • 2x faster inference and 60% lower costs vs. alternatives through research-optimized infrastructure
  • Largest open-source model ecosystem (200+ models) with native support for Llama, Mistral, Qwen, DeepSeek
  • True vendor independence—all models are open-source, fine-tuning output owned by user, can deploy anywhere
  • Enterprise-grade compliance: SOC 2 Type II, HIPAA-compliant with dedicated endpoints and reserved capacity
  • Cutting-edge research shipping to production—FlashAttention, speculative decoding, and custom kernels for measurable speedups

Cons

  • Token-based pricing complexity—variable rates per model make budget prediction difficult; no fixed-price tier for unpredictable workloads
  • Smaller ecosystem of integrations vs. OpenAI/Claude; LangChain/LlamaIndex support requires additional setup
  • Developer overhead—users must select, benchmark, and integrate models themselves; no opinionated defaults for non-technical teams
  • Learning curve for fine-tuning and GPU cluster management compared to fully managed services

Frequently Asked Questions

What is Together AI?

Together AI is an AI-native cloud platform that provides full-stack infrastructure for training, fine-tuning, and serving open-source models. It hosts 200+ models including Llama 4, DeepSeek R1, Qwen 3, and Mistral, and also supports audio TTS/STT, image generation, and video generation. The platform serves over 100 billion tokens daily and targets ML engineers, AI research teams, and enterprise AI teams.

How much does Together AI cost?

Together AI uses pay-as-you-go pricing with no subscription plans. Serverless inference ranges from $0.03 per million input tokens (small models) to $4.50 per million output tokens (large models like DeepSeek R1). Fine-tuning starts at $0.48 per million training tokens with a $4 per-job minimum. GPU clusters start at $3.99 per GPU-hour for H100 on-demand, with reserved pricing from $3.09 per hour.

Is Together AI free to use?

Together AI provides starter API credits to new accounts, which cover initial testing at no charge. There is no recurring free monthly allowance once those credits are used. After that, all usage is billed at pay-per-token or pay-per-GPU-hour rates with no minimum monthly spend.

What models does Together AI support?

Together AI hosts 200+ open-source models including Meta Llama 4 Maverick and Scout, DeepSeek R1 and V3, Qwen 3 and 3.5, Mistral, Gemma, and multimodal models for audio (Orpheus TTS, Whisper STT), image generation (FLUX.1, Google Imagen 4.0 Ultra), and video generation (Seedance 2.0, Google Veo 3.0).

Can I fine-tune models on Together AI?

Yes. Together AI supports full fine-tuning via SFT, DPO, and LoRA for models from 8B to 405B parameters, including Llama, Mistral, and Qwen variants. Users own their fine-tuned weights and can export them or deploy via dedicated endpoints. A cost estimation endpoint is available before committing to a training run.

What are the main features of Together AI?

Key capabilities include serverless inference across 200+ models, dedicated single-tenant GPU endpoints, on-demand GPU cluster rental (H100/H200/B200), fine-tuning with SFT/DPO/LoRA, batch processing at up to 50% discount, real-time audio TTS and STT, image generation (FLUX, Imagen), video generation (Seedance 2.0, Veo 3.0), managed storage, and Python SDK v2.0.

What compliance certifications does Together AI hold?

Together AI holds SOC 2 Type II, HIPAA, and ISO 27001:2022 certifications. HIPAA compliance requires using a dedicated endpoint rather than shared serverless infrastructure. The platform supports data processing agreements for enterprise customers, and data stays within user-selected regions.

Who is Together AI best for?

Together AI is best suited for ML engineers, AI research teams, and startup engineering teams that need cost-efficient inference, fine-tuning, or GPU clusters for open-source model development. It is less suited for teams needing a no-code or managed AI product, or those requiring exclusive compatibility with proprietary model APIs.

More AI Tools on HokAI

Visit Together AI Official Website