Best AI Text to Speech in 2026: What an Hour of Audio Really Costs
AI text to speech converts written text into spoken audio through an API or an editor. Vendors price it per 1,000 characters, per minute or per audio token. At 825 characters a minute, published list prices on 7 October 2026 work out to roughly $0.50 to $4 per hour of audio, depending on the model and vendor.
The short version
An hour of AI speech costs about $0.50 to $4 across ElevenLabs, Gemini, Murf, Deepgram and KugelAudio. Murf Falcon and Gemini Flash-Lite are cheapest, ElevenLabs is the broadest, and Gemini prices double on 1 January 2027. Free tiers all carry a cap, watermark or data condition.
The same hour of finished speech costs between about $0.50 and $4 depending on which text-to-speech service you pick, an eight-fold gap, and the cheapest option is not the one most guides put first. We read the live price pages of seven services on 7 October 2026 and converted every rate to one unit: dollars per hour of audio.
That unit matters because vendors quote three different things: per 1,000 characters, per minute of audio, or per million audio tokens. A buyer comparing a $0.03 rate with a €0.07 rate cannot tell which is cheaper without doing the arithmetic. We did it once, show the working, and mark where the numbers are our assumption and not the vendor's.
We did not generate audio for this guide, so it says nothing about which voice sounds best. It answers the question that comes before listening: what each option costs, what the free tier really gives you, and who each one is built for.
What does AI text to speech cost per hour of audio?
Short answer: Murf's Falcon model and Google's Gemini Flash-Lite TTS are the cheapest per hour, ElevenLabs and KugelAudio's premium model are the dearest, and Deepgram sits in the middle. All rates below are the vendors' published list prices for API use, checked on 7 October 2026.
| Option | Published rate | Per audio hour | Note |
|---|---|---|---|
| Murf Falcon | $0.01 per 1,000 characters | $0.50 | Built for voice agents |
| Gemini 3.8 Flash-Lite TTS | $6 per 1M audio tokens | $0.54 | Doubles on 1 Jan 2027 |
| Deepgram Aura-1 | $0.015 per 1,000 characters | $0.74 | Pay-as-you-go rate |
| Gemini 3.8 Flash TTS | $9 per 1M audio tokens | $0.81 | Doubles on 1 Jan 2027 |
| Murf Gen | $0.03 per 1,000 characters | $1.49 | Studio-quality model |
| Deepgram Aura-2 | $0.030 per 1,000 characters | $1.49 | Growth plan is $0.027 |
| ElevenLabs Flash / Turbo | $0.04 per 1,000 characters | $1.98 | 40,000 character limit |
| Deepgram Flux TTS | $0.0450 per 1,000 characters | $2.23 | Pay-as-you-go rate |
| KugelAudio kugel-2-turbo | €0.035 per minute | €2.10 | Hosted in the EU |
| ElevenLabs v4, v3, v2 | $0.08 per 1,000 characters | $3.96 | v4 is 72% off to 12 Oct |
| KugelAudio kugel-3 | €0.07 per minute | €4.20 | Premium model |
Read the table as a price ladder, not a quality ranking. The cheap end is aimed at phone bots and apps that speak in real time. The dear end sells expressive delivery and, in ElevenLabs' case, the widest set of voices and tools around the model.
The Gemini rows are the ones to watch. Google lists $0.00225 per 10 seconds of audio for the Flash model "through December 31, 2026" and $0.0045 from 1 January 2027, so the table shows today's rate and the bill you will have by next spring is double. If you build on it now, price your product at the 2027 figure. Google's own pricing page is the source for every Gemini number here, and its Google AI Studio entry in our directory covers how developers reach the models.
How we turned five price formats into one number
Per-hour cost needs an assumption about how many characters make an hour of speech. We used 825 characters per minute, the figure KugelAudio states on its own pricing page, which gives 49,500 characters per hour. A slow narrator will use fewer characters per hour and a fast one more, so treat each result as roughly right, not exact.
The working for each row:
- Per-1,000-character rates: multiply by 49.5. For example, Deepgram Aura-2 at $0.030 gives $1.485, shown as $1.49.
- Per-minute rates: multiply by 60. KugelAudio's kugel-3 at €0.07 per minute gives €4.20.
- Gemini audio tokens: Google states 25 tokens per second of audio, so an hour is 90,000 output tokens. At $9 per million that is $0.81. The text you send in adds about a cent more per hour, which we left out.
We did not convert euros to dollars, because the exchange rate moves and KugelAudio bills in euros before VAT. Rates also exclude tax, which ElevenLabs says on its own page.
One more caveat. These are API list prices. Consumer plans from ElevenLabs, Descript and Fliki bundle a monthly allowance of credits instead, and the per-hour figure depends on how much of it you use.
How to choose a text to speech service: the criteria
Decide on these five questions before you compare a single voice, because each one removes options.
- Pricing unit: characters, minutes or tokens, and whether the rate you were quoted is the real bill.
- Request limits: the maximum characters per call decides whether long text must be split.
- Licence: whether the plan you can afford allows commercial use.
- Data terms: whether your text may be used for training on that tier.
- Where it runs: an API for developers or an editor for creators.
The rest of this guide applies those five questions to each service in turn.
Where ElevenLabs wins
ElevenLabs is the broadest option and the most expensive per hour, and the range of choices inside it is the point. Its API page lists seven speech models at different price points, from $0.04 per 1,000 characters for the Flash and Turbo model up to $0.08 for v4, v3 and v2 Multilingual.
The models differ in ways that decide the bill. According to ElevenLabs, v4 supports 90+ languages and audio tags for fine control of delivery, v3 supports 70+ languages with a 5,000 character limit per request and multi-speaker dialogue, and the older v2 Multilingual covers 29 languages with a 10,000 character limit. Flash and Turbo cover 32 languages with a 40,000 character limit and a stated latency of about 75 milliseconds.
That 40,000 character ceiling is the practical detail for long narration. A chapter of an audiobook fits in one call on Flash, while v3 needs the text cut into pieces of 5,000 characters or fewer.
The consumer plans are a separate ladder, shown below as of 7 October 2026. A commercial licence starts at Starter, so the free plan is for trying the voices, not for publishing.
| Plan | Monthly price | Credits per month |
|---|---|---|
| Free | $0 | 10,000 |
| Starter | $6 (first month $1) | 30,000 |
| Creator | $22 (first month $11) | 121,000 |
| Pro | $99 | 600,000 |
ElevenLabs also sells a lot beyond speech: speech to text, sound effects, music, dubbing and voice changing, all priced on the same page. If you want one vendor for a whole audio workflow, that breadth is the reason to pay more. Our profile of ElevenLabs tracks the plans, and the voice assistants guide covers the conversational side of the same products.
Murf and Deepgram: priced for apps, not for hobbyists
Murf and Deepgram both list per-character API rates low enough to run a phone bot at scale, and both give new users free credit. They suit a developer, not someone who wants to paste a script into a web page and download an mp3.
Murf's API pricing page shows two speech models, compared with Deepgram's three in the table below. Falcon is described as optimised for conversational AI with time to first audio under 100 milliseconds. Gen is aimed at studio-quality voiceovers.
| Model | Per 1,000 characters | Positioning |
|---|---|---|
| Murf Falcon | $0.01 | Real-time voice agents |
| Murf Gen | $0.03 | Studio-quality voiceover |
| Deepgram Aura-1 | $0.015 | Lowest-cost Deepgram voice |
| Deepgram Aura-2 | $0.030 | Mid-range Deepgram voice |
| Deepgram Flux TTS | $0.0450 | Premium Deepgram voice |
Murf's free trial gives $10 of credit every month, with 150+ voices across 35+ languages and 5 concurrent requests. Its page lists SOC 2 Type II, ISO 27001, HIPAA and GDPR.
Deepgram started as a speech recognition company, and its rates above are the pay-as-you-go column. The Growth plan takes 10% off each. New accounts get $200 of free credit with no card required, and until 31 December 2026 Deepgram matches Flux TTS spend with credits up to $500.
The two are close enough on price that the deciding factors are the voices and the surrounding product. Deepgram bundles speech recognition and a voice-agent pipeline, which is useful if you are building the whole call. Murf also has a studio editor for non-developers, but we only captured its API prices, so we cannot give a plan price for it here. Read the Murf AI and Deepgram entries, and the Vapi and Retell AI pages if what you need is the agent layer on top of a voice.
KugelAudio: the pick when the data has to stay in Europe
KugelAudio is the only service here that leads with where the servers are. Its homepage says the voices are developed and hosted in Europe and are fully GDPR compliant, and the pricing page says an EU Sovereign Cloud is included on every plan for EU and EEA customers without an enterprise contract.
Prices are per audio minute, billed by the second: kugel-2-turbo is €0.035 a minute, about €42.42 per million characters on KugelAudio's own conversion, and kugel-3 is €0.07 a minute. Both list voice cloning, speed control, word-level timestamps and an integrated normaliser, and the site covers 39 languages.
Its own page claims the fastest median time to first audio among the models it tested, at 91 milliseconds against 126 for ElevenLabs Flash v2.5. That is the vendor's measurement of its competitors, which is not independent, so treat it as a claim to test with your own traffic.
The normaliser is the feature to try. KugelAudio shows IBANs, postcodes and phone numbers read digit by digit as its selling point, and those are exactly the strings that trip a generic voice in a banking or clinic call. Our KugelAudio profile has the details. If you would rather call open models through a gateway, Together AI and Replicate both host speech models, though we did not price them here.
Gemini TTS: cheapest from a big vendor, with a catch
Gemini's speech models are the lowest per-hour price from a large provider, and the free tier has a condition most people skip. Google's pricing table lists "Used to improve our products" as Yes on the free tier and No on the paid tier.
In practice that means text you send on the free tier may be used by Google to improve its products. For a hobby project that is fine. For a script that contains customer names, it is not, and the paid tier is the one to use.
Google lists four ways to buy the same model. Standard is the headline rate, Batch and Flex are half of it for work that can wait, and Priority is $16.20 per million audio tokens (until the end of 2026) for faster service. For narration that is rendered overnight, the Batch rate works out to about $0.41 per audio hour on the Flash model today.
The trade is reach. You get a model and an API, not a studio, so there is no timeline editor, no pronunciation dictionary UI and no voice marketplace. Gemini has a profile in the directory, next to Google's wider product set. OpenAI also sells speech models, which we did not price here.
Fliki and Descript: when the voice is part of a video
Fliki and Descript are for people who make videos, not for people who need an API. Speech is one feature inside a tool that also edits and exports video, and both price in credits and media hours, not characters.
Fliki and Descript compare like this on the plan features we could read:
| Plan | Allowance | Voice and video features |
|---|---|---|
| Fliki Free | 3 credits a month | 300 voices, 720p, watermark |
| Fliki Standard | 180 credits | 1,000 voices, 1080p, voice cloning |
| Fliki Premium | 600 credits | 2,000+ voices, 40 minutes, API |
| Descript Hobbyist | 10 media hours | AI speech with voice clones |
| Descript Business | 40 media hours | Dubbing in 30+ languages |
Fliki's estimator says a five-minute video with an AI voiceover costs about 2.5 credits. We could not capture the dollar prices on Fliki's page, so we leave them out.
Descript prices per person per month and shows two figures for each plan, which we read as annual against monthly billing. Hobbyist starts at $16 and Business at $50 on the lower figure.
If your starting point is a script and a finished video, start there.
Avatar tools such as HeyGen and Synthesia bundle a voice too, and our AI video generators category and content creation category list the rest. For the whole workflow, see how to create AI video without a production team. If it is a pure voice file, the API services above cost less per hour and drop the video features you do not need. Compare the video side in our video generators guide, the YouTube video guide and the video editing guide, or browse Fliki, Descript, VEED and CapCut directly.
The shortlist: best AI text to speech picks by job
Match the job to the pricing model first and the voice second. These are our picks from the published prices, not from listening tests.
| The job | Our pick | Why |
|---|---|---|
| Phone voice agent | Murf Falcon or Deepgram Aura | Low-latency models at the low end of the price ladder |
| Long expressive narration | ElevenLabs Flash, or v3 and v4 | Flash takes 40,000 characters per call |
| Cheap overnight batch | Gemini Flash-Lite, Batch rate | Paid tier keeps your text out of training |
| EU data residency | KugelAudio | Hosted in Europe, dearest per hour |
| Explainer videos | Fliki or Descript | An editor around the voice |
Still not sure? Answer a few questions in Smart Match and it narrows the shortlist from our directory of AI tools.
The wider list sits in our AI voice assistants category, where Play.ht, Resemble AI and Lalals also have profiles. We left Resemble out of the price table because its pricing page leads with deepfake detection plans, so we could not find a speech rate to compare.
What does "free and unlimited" really get you?
Nothing in this price survey is unlimited. Every free tier we read has a cap, a watermark, a licence limit or a data condition, and "free unlimited" in a search result usually means one of those is hidden.
- ElevenLabs Free: 10,000 credits a month and no commercial licence until Starter.
- Murf: $10 of credit every month on the API trial.
- Deepgram: $200 of one-time credit, no card.
- Gemini free tier: no charge, but your inputs may be used to improve Google's products.
- Fliki Free: 3 credits a month and a watermark on 720p video.
For a test, any of these is enough. For a podcast or a product, price the paid tier.
How is your text handled?
Check the data terms before the price. The only service in our set that states a training condition by tier on its price page is Google, as covered above. Descript lists an opt-out of training only on its Enterprise plan, so read each vendor's terms for the plan you use.
Murf names its compliance badges and KugelAudio names its EU hosting, but both are claims on marketing pages. If your text is regulated, ask for the audit report or the data processing agreement before you send a line of it.
The case against buying on price per hour
The strongest argument against this guide's ranking is that a price per hour of generated audio is not a price per hour of finished audio. If a cheap model needs three takes to get a sentence right and a dear one needs one, the dear one is cheaper.
That is a fair objection and we cannot answer it with data, because we did not generate audio or count retakes. What the prices do tell you is the ceiling on the damage. An hour of narration costs between $0.50 and $4 across the whole set, so even a 4x retake rate on the cheapest option stays below the price of one pass on the dearest. For a voice agent that speaks millions of characters, the maths is different, which is why we split the picks by job above.
The practical test is small. Generate the same 500 words on two services, listen, and count how many sentences you would redo. It takes ten minutes and costs well under a dollar, even at ElevenLabs' $0.08 rate.
What would change this
Three things would move the ranking:
- Google's scheduled 2027 rise takes Flash-Lite above Deepgram Aura-1 per hour, so the cheapest big-vendor option stops being cheapest.
- If ElevenLabs makes its v4 launch price permanent, its per-hour figure falls to about a quarter.
- If Murf publishes its studio plan prices, the picture for non-developers changes.
Two of the prices above carry end dates (ElevenLabs v4 on 12 October, Gemini on 31 December), so re-check the vendor page before you commit to a budget. The dated source list below shows what we read and when. If you are also choosing the language model that writes the script, the model leaderboard shows what each one costs per token.
What we could not verify
Be clear about the limits of this guide:
- We did not listen to or generate audio, so voice quality is not ranked.
- Fliki's dollar prices and Murf's studio plan prices were not available on the pages we read.
- The per-hour figures use an assumed 825 characters a minute.
- KugelAudio's latency comparison is its own claim, not an independent test.
- We could not load Play.ht's site on 7 October, so it has no price here.
Frequently asked questions
What is the cheapest AI text to speech API?
On published list prices checked on 7 October 2026, Murf Falcon at $0.01 per 1,000 characters and Gemini 3.8 Flash-Lite TTS at $6 per million audio tokens come to about $0.50 and $0.54 per audio hour. Google doubles its Gemini rates on 1 January 2027, so the Gemini figure will not hold.
Is there a free unlimited AI text to speech service?
Not among the services we checked. ElevenLabs Free gives 10,000 credits a month without a commercial licence, Murf gives $10 of API credit monthly, Deepgram gives $200 once, and Gemini has a free tier where inputs may be used to improve Google products. Each has a cap, a licence limit or a data condition.
How much does ElevenLabs cost per hour of audio?
ElevenLabs lists $0.04 per 1,000 characters for Flash and Turbo and $0.08 for v4, v3 and v2 Multilingual on its API page. At 825 characters a minute that is roughly $1.98 and $3.96 per hour. Subscription plans instead bundle credits, starting at $6 a month for Starter.
Is Gemini text to speech free?
Google lists a free tier for its Gemini TTS models, but its pricing table says free-tier content is used to improve its products and paid-tier content is not. Paid use is $9 per million audio tokens for Gemini 3.8 Flash TTS through 31 December 2026, then $18.
Which text to speech service keeps data in Europe?
KugelAudio states its voices are developed and hosted in Europe and that an EU Sovereign Cloud is included on every plan for EU and EEA customers. Rates are 0.035 euros per minute for kugel-2-turbo and 0.07 euros for kugel-3. Ask for its data processing terms before sending regulated text.
Covered in this guide
- Deepgram: Real-time voice AI APIs for developers: speech-to-text under 300ms, TTS, and voice agents. Used by 200,000+ developers.
- ElevenLabs: The leading AI voice platform for text-to-speech, voice cloning, and conversational AI agents
- KugelAudio: European TTS API with fast first-audio latency across dozens of languages, GDPR-compliant EU hosting with zero data retention, and a 78% human-preference win rate over ElevenLabs on German-language tests. Billed per minute, no card required to start.
- Murf AI: Ultra-realistic AI Voice Generator with fastest text-to-speech API for voice agents and creators
- CapCut: AI-powered video editor for creating trending content across all platforms
- Descript: AI-powered video & audio editing as easy as editing text
- Fliki: Transform text into professional videos with AI voiceovers in minutes
- Gemini: Google's multimodal AI assistant and model family, sold as a free app, tiered subscriptions and a pay-per-token API
- Google AI Studio: Free web-based IDE for building, testing, and deploying generative AI applications with Google's Gemini models
- HeyGen: Fast, simple AI video generator that creates professional videos with realistic avatars in 175+ languages.
- Lalals: All-in-one AI audio platform for music creators: 1,000+ voice models, stem splitting, voice cloning, AI mastering, music/sound generation, and audio cleanup (de-noise, de-reverb, de-echo): credit-based subscription pricing.
- OpenAI: OpenAI builds ChatGPT, the GPT-6 family (Astra, Sol, Luna), Codex and the OpenAI API. It reports 1.2B weekly ChatGPT users and an $852B valuation after its March 2026 round.
- Play.ht: [DISCONTINUED] Play.ht was an AI text-to-speech platform with 900+ voices in 142 languages. Meta acquired the company in 2025 and shut it down completely by year's end. See TechCrunch and Meta's official announcement for proof.
- Resemble AI: Generative voice AI and deepfake detection for enterprise trust
- Retell AI: Retell AI builds and deploys realistic AI phone agents with 600ms latency, handling 50M+ monthly calls, and HIPAA/SOC 2 compliance. Rated 4.8/5 on G2 by 1,414 teams.
- Synthesia: AI video platform transforming text into professional, multilingual videos with realistic avatars
- Vapi: Voice AI platform where developers bring their own LLM, speech-to-text and telephony; Amazon Ring picked it over 40 rivals in 2026.
- VEED: Browser-based video editor with AI subtitles, avatars, voice cloning, and text-to-video: no download needed.
Sources
Still deciding?
This guide covers a handful of options. Smart Match checks every listing in the directory against how you actually work and what you can spend, then hands you the shortlist and the reason behind each pick.
Start Smart MatchRelated guides
- 41 of 100 HokAI Guides Carry a Screenshot. All 106 Still Load.Inside HokAIHow the directory works, verified in the open
- AdsMaker.ai vs Creatify: Creatify Tripled Its Price, AdsMaker Didn'tComparisonHead-to-head, with a verdict
- The AI 80s Photo Trend: How to Make Yours (and Turn It Into a Video) in 2026AnalysisWhat changed and who it affects
- Best Agentic AI in 2026: Pick the Job, Not the GeneralistBuyer's guideHow to pick, across a category
- Best AI Assistant Apps for Android in 2026, Now That Google Assistant Is DeadBuyer's guideHow to pick, across a category
- Best AI Chatbots in 2026: Pick by the Job, Not the LeaderboardUpdatedRechecked against current sources