Last updated: 2026-08-19
Deepgram's Nova-3 model posts a 5.26% word error rate on English speech, streamed in real time via WebSocket. The company's APIs also cover text-to-speech and full Voice Agent orchestration for developers, with SOC 2 Type II, HIPAA, and GDPR compliance for regulated transcription workloads.
About Deepgram
Deepgram is a voice AI platform built by Deepgram Inc. (San Francisco, founded 2015), providing APIs for speech-to-text, text-to-speech, and conversational voice agents used across customer support, healthcare, media, and conversational AI. The company has raised $246 million in total funding and reached a $1.3 billion valuation as of January 2026. At the core of the platform are neural speech models trained end-to-end for speed and accuracy. The flagship Nova-3 model streams transcripts in real time via WebSocket, faster than most competitors. The Flux model, launched in 2024, is designed for conversational AI specifically: it handles turn detection and natural interruptions without additional tooling. All models support 45+ languages and bill per second rather than per minute, so teams pay only for audio actually processed. Deepgram targets backend engineers and product teams building real-time voice features. Common use cases include call center analytics (detecting sentiment and topics in live calls), voice agent pipelines that chain STT with an LLM and TTS, live captioning, and accessibility tooling. The add-on intelligence features (speaker diarization, smart formatting, entity redaction) make it viable for healthcare and legal transcription workflows. Pricing is usage-based and billed per second rather than per minute. The Pay-As-You-Go tier starts with free credits and scales through a Growth plan requiring an annual prepayment, up to a self-hosted Enterprise option for teams with full data-sovereignty requirements. The platform is accessible via cloud API and web dashboard; no native desktop or mobile app exists. In January 2026, Deepgram closed a $130 million Series C round and simultaneously acquired a Y Combinator AI startup. SDKs are available in Python, JavaScript, Go, and .NET, with REST and WebSocket endpoints covering all products. GitHub: github.com/deepgram.
Pricing
Pay-As-You-Go: $200 free credit to start, then $0.0077/min Nova-3 mono, $0.0092/min multilingual, $0.08/min Voice Agent. Growth: $4,000+ annual prepayment (~20% discount, ~$0.0065/min). Enterprise: $15,000+/year with custom pricing and self-hosted option. Billed per second.
Key Features
- Nova-3 Real-Time STT: Delivers transcripts via WebSocket in under 300ms, achieving 5.26% WER on English and supporting 45+ languages, with a 34% batch WER reduction released in March 2026.
- Voice Agent API: Bundles STT, TTS, and conversation orchestration into one streaming endpoint, with built-in barge-in detection, turn-taking, and mid-conversation function calling.
- Per-Second Billing: Charges are calculated per second of audio processed (not rounded to the nearest minute), reducing costs by 10-50% on short clips compared to per-minute competitors.
- Speaker Diarization: Automatically identifies and labels individual speakers in multi-party audio without requiring the number of participants to be specified in advance.
- Audio Intelligence Add-ons: Sentiment analysis, topic detection, smart formatting, and PII entity redaction are available as separately billed add-ons layered on top of base transcription in a single API call.
- Native MCP Server: Official Model Context Protocol server (April 2026) connects Claude Code, Cursor, and Windsurf directly to Deepgram's API for agentic transcription and TTS without leaving the dev environment.
Pros
- Real-time streaming latency runs roughly 3x faster than AWS Transcribe's comparable endpoint, at a similar per-minute price point.
- Per-second billing means a 37-second audio clip costs $0.00474 exactly, not rounded to a full minute, saving 30-40% vs. per-minute billing competitors.
- Nova-3 posts a lower word error rate on English than Google's Chirp 2 model, at less than half Chirp 2's per-minute price for production transcription work.
Cons
- Purely an API product with no consumer app: non-technical users cannot use Deepgram directly without a developer building an interface for them.
- Intelligence add-ons (speaker diarization, sentiment, entity redaction) each carry separate per-minute charges, making complex pipelines 3-5x more expensive than base transcription.
- Concurrent connection limits on Pay-As-You-Go and Growth plans can cause dropped WebSocket connections during sudden traffic spikes in high-scale live event deployments.
Data Handling
- Training-data policy
- Does not train on customer data by default. Opt-in only via the Model Improvement Partnership Program.
- Data retention
- Zero retention
- Compliance
- SOC 2 Type II · HIPAA · GDPR · PCI DSS · CCPA
Frequently Asked Questions
What does Deepgram actually cost?
Pay-As-You-Go starts with $200 in free API credits, then Nova-3 transcription runs $0.0077 per minute (monolingual) or $0.0092 per minute (multilingual), with the bundled Voice Agent API at $0.08 per minute. The Growth tier requires a $4,000+ annual prepayment for roughly 20% off, and Enterprise starts at $15,000 per year with a self-hosted option. All usage bills per second rather than rounding up to the next minute.
Does Deepgram have a free plan?
No permanent free plan exists. New Pay-As-You-Go accounts open with non-expiring API credit, good for roughly 430 hours of Nova-3 monolingual transcription. No subscription or minimum spend is attached, and usage past that credit bills at standard per-second rates.
What should you use instead of Deepgram?
AssemblyAI is the strongest swap when transcription needs deeper built-in audio intelligence models alongside it. Speechmatics does more for multilingual accuracy work. Rev AI appeals to teams that weigh the lowest per-minute rate above real-time streaming speed.
Deepgram or AssemblyAI: which should you pick?
Deepgram wins on raw transcription speed, streaming results in real time with lower latency than AssemblyAI's default endpoint, and undercuts it on list price for high-volume real-time use. AssemblyAI counters with a broader out-of-the-box audio intelligence suite (summarization, content moderation) that would otherwise require Deepgram's separately billed add-ons.
How long does it take to get going with Deepgram?
Minutes to a first transcript for most developers. Registration at deepgram.com releases free API credit without a card, and the dashboard issues a key straight away. Send a short audio file through any official SDK, or call the REST or WebSocket endpoint directly, following the quick start at developers.deepgram.com.
Top Alternatives
- AssemblyAI: AssemblyAI brings the broader built-in audio intelligence suite; Deepgram answers with faster real-time streaming and per-second billing.
- Speechmatics: Speechmatics is the call when multilingual accuracy across 55+ languages leads, while Deepgram holds lower-latency English-first streaming.
- Rev AI: Rev AI is cheapest per minute for batch transcription, whereas Deepgram exists for sub-second real-time streaming.
HokAI guides covering Deepgram
- What Murf AI Is Actually For: Murf AI now runs two products: a $19-$99 narration tool and a new $0.01-per-minute API for real-time voice agents. What each one is actually built for.
- What Skilly Is Actually For, and Which Skilly You Mean: Skilly is filed as an AI voice assistant next to Siri, but it is really two small products forked from Clicky in five days. Here is which one is worth using.
- Best AI Voice Assistants in 2026: A Buyer's Guide to Synthesis, Transcription and Agents: AI voice assistants split into three buying decisions: synthesis, transcription and agent orchestration. Compare ElevenLabs, Deepgram and Retell AI for 2026.
- Murf AI vs Skilly: Why This Comparison Is Asking the Wrong Question: Murf AI makes voiceover audio and voice agents; Skilly is a $19/month Mac tutor that narrates Figma and Blender live. Here's which job needs which tool.