Last updated: 2026-10-11
Cartesia's Sonic-3.6, generally available since late August 2026, is a streaming text-to-speech model built on state space models rather than transformers. The San Francisco company, founded in 2023, pairs it with Ink-2 speech-to-text and Managed Agents for phone voice agents, all reached through a web playground and an API.
About Cartesia
Cartesia is a San Francisco voice AI company founded in 2023 by Karan Goel, Albert Gu, Arjun Desai and Brandon Yang, who met as PhD students in Stanford's AI lab and worked on state space models, an alternative to the transformer design behind most large language models. It raised a $27 million seed round led by Index Ventures in December 2024 and a $64 million Series A led by Kleiner Perkins in March 2025. In October 2025 its chief executive announced a further $100 million from Kleiner Perkins, Index Ventures, Lightspeed and NVIDIA alongside the launch of the Sonic-3 model.
Three products share one account. Sonic is the text-to-speech model, in its current release speaking 44 languages and accepts inline cues such as [laughter] plus pronunciation dictionaries for names and technical terms. Ink-2 is the speech-to-text model, built to signal when a caller starts and stops talking without a separate voice-activity detector. Managed Agents chains Ink-2, a language model you choose and Sonic into a phone-ready voice agent with tool calls, SIP trunking and built-in evaluations. Speech and transcription draw on one monthly credit pool, while agent minutes are paid from separate prepaid dollars.
The natural buyer is a developer or product team shipping a voice agent for support lines, outbound sales calls or candidate screening, where the delay before the first spoken word shapes how natural a call feels. Cartesia publishes official Python and JavaScript SDKs, a code-first Line SDK and an MCP server on GitHub, and the enterprise tier adds on-premises, private-cloud and OEM deployment. Teams that want to mix their own speech, language-model and telephony vendors can pair the voice with an orchestration layer such as Vapi or Retell AI, and the vendor says Sonic can also run inside your own cloud through Together AI or Baseten.
Weigh it against the field before committing. The wider audio catalog sits with ElevenLabs, cloning with published open weights with Fish Audio, a studio timeline editor with Murf AI, and watermarking plus deepfake detection with Resemble AI. Deepgram bills transcription by the second, while AssemblyAI specializes in speech understanding. Buyers who need transcription in more languages than Ink-2 covers can look at Speechmatics, which lists 55-plus languages and on-premises deployment, or Rev AI. The vendor also says Ink-2's built-in noise handling removes the need for add-on filters such as Krisp. Our hour-of-audio price comparison costs out the speech providers, our voice assistants roundup is the companion read, and the voice and conversational tools category lists the rest.
Pricing
List prices in USD per month, read from the vendor's pricing page in October 2026: Free $0, Pro $5, Startup $49, Scale $299, Enterprise custom. 25 million and 8 million credits) plus prepaid dollars for voice agents equal to the plan price, with $1 on Free. One credit buys one character of speech, a minute of Sonic audio uses 750 to 800 credits and Ink-2 transcription uses 3 credits a second.
Divided out, the paid plans work out to about $50, $39 and $37 per million credits on Pro, Startup and Scale. Unused credits roll over up to twice the monthly allowance, and a plan change keeps what you already paid for. Character pricing can be compared on the ElevenLabs page, and per-minute agent rates are on the Vapi page, with Retell AI as a second reference.
| Tier | Monthly price | What it includes |
|---|---|---|
| Free | Free | 20,000 credits a month, about 27 minutes of speech, 2 concurrent speech requests and 8 transcription requests, 1 Cartesia phone number and 8 concurrent agent calls, $1 of prepaid voice-agent spend a month |
| Pro | $5/mo | 100,000 credits a month, about 133 minutes of speech, Commercial use licence, Instant voice cloning, 3 concurrent speech requests and 12 transcription requests, 3 phone numbers and 12 concurrent agent calls, $5 of prepaid voice-agent spend a month |
| Startup | $49/mo | 1.25 million credits a month, about 1,667 minutes of speech, Professional voice cloning with 2 slots, Organizations for team accounts, 5 concurrent speech requests and 20 transcription requests, 5 phone numbers and 20 concurrent agent calls, $49 of prepaid voice-agent spend a month |
| Scale | $299/mo | 8 million credits a month, about 10,667 minutes of speech, Priority support, Professional voice cloning with 4 slots, 15 concurrent speech requests and 60 transcription requests, 10 phone numbers and 60 concurrent agent calls, $299 of prepaid voice-agent spend a month |
| Enterprise | Custom | Volume pricing and custom concurrency limits, Data processing and business associate agreements, Single sign-on and a shared Slack channel, Security questionnaires, On-premises, private-cloud and OEM deployment by contract |
| Voice agents (usage) | $0.06 per minute | Call duration billed per minute, Cartesia-provided phone numbers $0.014 a minute, Language-model usage free for dashboard-built agents for a limited time, Evaluations free for a limited time |
Key Features
- Streaming text-to-speech: The current Sonic release speaks 44 languages and the vendor reports first audio in under 90 milliseconds, the delay voice-agent teams budget around.
- Expressive and controllable delivery: Inline tags such as [laughter] add non-verbal sounds, and custom pronunciation dictionaries take phonetic spellings for proper nouns and domain terms such as drug names.
- Instant and professional cloning: Instant cloning works from a short clip and Professional Voice Cloning trains on 15 to 30 minutes of audio, with training itself costing no credits.
- Turn-aware transcription: Ink-2 emits turn start, turn end and an eager-end signal from the model itself, so an agent can begin its reply before the caller has fully finished.
- Managed voice agents: Managed Agents wires transcription, a language model of your choice and Sonic into one runtime that handles turn-taking, tool calls, telephony and call analytics.
- Private deployment options: Enterprise contracts allow on-premises installs including air-gapped sites, a private cloud on AWS, Google Cloud or Azure, or OEM licensing inside your own product.
- Developer tooling: Official Python and JavaScript SDKs, a Line SDK for code-first agents and an MCP server are published on the vendor's GitHub, with the SDKs under Apache-2.0.
Pros
- Speech output, transcription and agent calls each get their own concurrency pool, so a surge in speech requests cannot crowd out live calls on the same account.
- Dated model snapshots never change once released, which lets a production team freeze behavior while it runs its own evaluations before moving on.
- G2 reviewers rate Cartesia 4.3 out of 5 across 43 reviews, a thinner sample than the 1,253 reviews behind ElevenLabs on the same site.
- The Trust Center lists a SOC 2 Type II report, a penetration-test report and a HIPAA document, and the Data Protection Addendum is public, which shortens a security review.
Cons
- The vendor's site markets Sonic as ranked first in the Speech Arena, yet Artificial Analysis's provider-voice leaderboard, read on 11 October 2026, shows Sonic 3.6 at an Elo of 1284 against 1323 for ElevenLabs' Eleven v4, so treat the top-spot line as out of date.
- Under the standard terms, inputs, outputs and usage may be used to train Cartesia's models unless an agreement says otherwise, an opt-out form only stops future use, and zero data retention is a setting tied to the addendum rather than a default.
- Ink-2 transcribes only English, Spanish, French, Hindi and Japanese, so a voice agent that must listen in German, Korean or Arabic cannot use Cartesia's own speech-to-text even though Sonic can speak those languages.
- The vendor's pages disagree in places: the pricing table marks instant cloning as unavailable on Free while the Sonic FAQ lists it there, and the docs cite a 10-second sample while the FAQ says under a minute, so confirm with support before you build on either.
Data Handling
- Training-data policy
- Standard terms let Cartesia use inputs, outputs and usage to train its models unless an agreement says otherwise. A form lets you opt selected content categories out of future training. A zero data retention setting stops storage of audio, transcripts and outputs, except voice samples supplied for cloning.
- Compliance
- SOC 2 Type 2 · HIPAA · GDPR · PCI DSS 4.0.1 (vendor Trust Center)
Frequently Asked Questions
What does Cartesia actually cost?
Self-serve plans are $5 a month for Pro, $49 for Startup and $299 for Scale, and Enterprise is custom. Going past the monthly credit pool needs overage switched on, billed at $65, $45 and $38 per million credits on those three plans, so usage beyond roughly 775,000 credits a month costs less on Startup than on Pro with overage. Voice agents are a separate meter at $0.06 a minute, covered first by the plan's prepaid dollars (about 83 hours on Scale), and a Cartesia phone number adds $0.014 a minute.
Can you use Cartesia without paying?
Yes. The Free plan gives 20,000 credits a month, enough for about 27 minutes of speech or nearly two hours of transcription, with 2 concurrent speech requests. It has no commercial licence and no instant cloning, and the terms bar commercial use of outputs unless your tier allows it, so treat Free as a place to audition voices and measure latency rather than to ship.
What are Cartesia's closest competitors?
It depends on the layer you are buying. For speech output, ElevenLabs and Murf AI are the usual names, with Murf geared to slide and video voiceovers. For a whole call stack, Vapi and Retell AI orchestrate calls around a voice of your choice, and for transcription alone Deepgram and AssemblyAI are the specialists.
How does Cartesia compare to ElevenLabs in 2026?
Cartesia's own August 2026 listener test says 92 percent of US English votes preferred its latest Sonic over Eleven v3, but that test is vendor-run and the comparison model is older than ElevenLabs' current Eleven v4 line. Artificial Analysis's provider-voice leaderboard, which uses blind votes from outside listeners, listed the two Eleven v4 models above Cartesia's newest Sonic when we read it on 11 October 2026. Pick ElevenLabs if you want the broader product catalog and Cartesia if agent latency, on-premises deployment or one-vendor call handling decides the purchase.
How do you set up Cartesia?
The docs say you can audition voices in the online playground without writing code, and an account on the Free plan then gives you an API key for the Python or JavaScript SDK. Use the rolling model id from the docs to follow the latest stable release, or a dated snapshot id to freeze behavior while you test. For a phone agent, create a Managed Agent in the dashboard, choose a voice, language and language model, attach tools, then run test calls before attaching a number.
Top Alternatives
- ElevenLabs: Pick ElevenLabs for the wider catalog, music app included; pick Cartesia when streaming agent latency and one vendor for speech, transcription and agent runtime matter most.
- Deepgram: Pick Deepgram for pay-as-you-go, per-second billing on speech recognition; pick Cartesia when you want Sonic voices and cloning alongside transcription.
- Vapi: Pick Vapi to bring your own language model, transcription and telephony; pick Cartesia when you want its own Ink and Sonic models and a managed runtime from one vendor.
- Retell AI: Pick Retell AI for a no-code conversation-flow builder and outbound campaign tools; pick Cartesia when the voice models themselves are what you are buying.
- Fish Audio: Pick Fish Audio for a free development API and published open weights for research; pick Cartesia when SOC 2 Type II and HIPAA paperwork should already be on a Trust Center.