by Speechmatics Ltd.

Speechmatics review, pricing and verdict

Speech-to-text API from Cambridge UK covering 55+ languages. Medical model at 93% accuracy, on-prem deployment. Free 4 hrs/month, on-demand from $2.75/hr.

  • ai voice assistants
  • Web
  • Windows
  • Mac
  • iOS
  • Android
checked

Last updated: 2026-08-19

Speechmatics is a Cambridge, UK speech-to-text API returning real-time partial transcripts in under 250ms. It differentiates through a specialized medical transcription model and support for four deployment modes: managed cloud, self-hosted, virtual appliance, and on-device edge processing. SOC 2 Type II, HIPAA, and ISO 27001 certified.

About Speechmatics

Speechmatics is a speech AI platform built by Speechmatics Ltd., founded in 2006 in Cambridge, UK by Dr. Tony Robinson, a pioneer in applying neural networks to speech recognition. With $90.6 million in total funding including a $62 million Series B in 2022, the company provides APIs for speech-to-text, real-time transcription, and voice AI agents, used across enterprise customers in finance, healthcare, broadcasting, and government. The platform differentiates on three axes: the widest language and dialect coverage in the market, including regional accents that generic models struggle with; a specialised medical transcription model; and deployment flexibility across cloud, self-hosted, virtual-appliance, and edge environments for air-gapped or battery-sensitive settings. In September 2025, Speechmatics launched a specialized Medical Speech-to-Text model for clinical documentation, cutting keyword error rates well below the nearest competing system. An Enhanced model tier targets maximum accuracy on noisy or accented audio, while a Standard tier is optimized for speed on batch work. Enterprise plans carry custom pricing and include dedicated support, service-level agreements, and optional HIPAA-compliant processing on top of the standard usage-based rates. Speechmatics is best suited to regulated-industry teams and enterprises where data sovereignty, accent coverage, or medical-grade accuracy are firm requirements. It is not cost-competitive for high-volume English-only workloads, where Deepgram Nova-3 at $0.46/hr or Rev.ai Reverb at $0.18/hr are more affordable alternatives.

Pricing

Free: 4 hours/month (2 hrs Enhanced + 2 hrs Standard). Volume discount: 20% off usage above 500 hrs/month; additional discounts at 24k+ hrs/year. Enterprise: custom pricing with SLAs and HIPAA mode. Self-hosted or on-premises deployments incur separate infrastructure costs beyond the standard hourly rate.

Key Features

  • Real-Time Streaming STT: Sub-250ms latency for partial transcripts via WebSocket, enabling live captioning and voice agents without network round-trip delays.
  • Medical Model for Clinical Use: Trained on clinical conversations, the model reaches clinical-grade real-world accuracy with a keyword error rate 50% lower than the nearest competing system, including 96% recall on medical terminology.
  • Speaker Diarization Included: Automatic speaker identification and labeling ships as a core feature, not a paid add-on, identifying individual speakers without requiring a speaker count in advance.
  • Four Deployment Modes: Runs as a managed cloud API, a self-hosted Docker container, a virtual appliance, or a local on-device runtime: the only vendor supporting all four modes without requiring an internet connection for edge processing.
  • 55+ Languages with Dialects: Covers regional and dialect variants across Scandinavian, Germanic, Romance, Asian, and Indian languages, including bilingual code-switching models such as Arabic-English.
  • On-Device Model for Privacy: An April 2026 integration with Adobe Premiere Pro runs a local model within 5% of cloud accuracy on hardware such as Mac M5, NVIDIA RTX, and Intel or AMD CPUs, with zero data leaving the device.
  • Domain Adaptation and Custom Dictionaries: Fine-tune models on customer audio and terminology to improve accuracy on domain-specific jargon in legal, medical, financial, and technical use cases without separate retraining infrastructure.

Pros

  • Speaker diarization is bundled into every plan at no extra cost, unlike Deepgram and AssemblyAI, which both charge a separate per-minute add-on fee for the same capability.
  • The on-device model is a standout differentiator: no other major STT vendor ships a laptop-runnable model with near-cloud accuracy, which matters for teams handling audio that legally cannot leave the device.
  • The medical model's clinical-conversation training gives it deeper healthcare specialization than any general-purpose STT vendor, which is why regulated healthcare teams pick it over cheaper general competitors.

Cons

  • Speechmatics costs six to fifteen times more per minute than Deepgram ($0.0077/min) or Rev.ai ($0.003/min) for basic English batch transcription, which prices out cost-sensitive startups doing high-volume work.
  • Accuracy varies across the language portfolio: some languages such as Turkish and Arabic still trail English performance, so teams should test and tune before relying on them in production.
  • The free tier caps out at 4 hours a month, tight next to competitors offering 25+ free hours, so evaluating Speechmatics at scale can mean hitting the paywall quickly.

Data Handling

Training-data policy
Does not train on customer audio by default. Medical model trained on consented clinical conversations with privacy guarantees.
Data retention
30 days
Compliance
SOC 2 Type II · ISO/IEC 27001:2022 · HIPAA · GDPR · PCI DSS

Frequently Asked Questions

What does Speechmatics actually cost?

Speechmatics costs $2.75 per hour on the Standard model and $3.75 per hour on Enhanced, billed on demand with no subscription commitment. Usage above 500 hours a month gets a 20% volume discount, with further discounts past 24,000 hours a year. Enterprise plans add custom SLAs and HIPAA-compliant processing at negotiated pricing.

What do you get on Speechmatics's free tier?

Four hours a month come free, split into 2 hours on the Enhanced model and 2 hours on Standard, with no payment details needed to start. The allowance resets each month rather than rolling over, so unused hours disappear. It covers evaluation comfortably but runs out fast on any production workload.

What should you use instead of Speechmatics?

Deepgram costs less for real-time English transcription, as long as on-premises or edge deployment isn't a requirement. Rev.ai undercuts Speechmatics further on basic English batch work but offers no multilingual or edge support. AssemblyAI makes more sense when sentiment analysis or summarization needs to be built in, rather than raw transcription alone.

Is Speechmatics better than Deepgram?

Deepgram is faster and cheaper for real-time English-only transcription, but Speechmatics wins on language breadth, deployment flexibility, and clinical-grade medical transcription that Deepgram does not offer. Teams needing on-premises, edge, or air-gapped deployment have no equivalent option on Deepgram. For cost-sensitive, English-only, high-volume pipelines, Deepgram remains the cheaper choice.

How long does it take to get going with Speechmatics?

Getting a first transcript running takes about as long as generating an API key: sign up at speechmatics.com, then call the REST endpoint for batch transcription or open a WebSocket connection for real-time streaming. Official SDKs cover JavaScript and Python, with additional guides for React Native and LiveKit voice agents, and full documentation lives at docs.speechmatics.com.

Top Alternatives

  • Deepgram: Sub-300ms real-time English transcription from $0.0077/min defines Deepgram; broader language and dialect coverage, a medical-grade model, and on-premises or edge deployment define Speechmatics.
  • Rev AI: Rock-bottom English batch pricing at $0.003/min is Rev AI's case; on-premises or edge deployment and a clinical-accuracy medical model Rev AI doesn't offer make the case for Speechmatics.
  • AssemblyAI: Built-in audio intelligence like sentiment analysis and summarization is AssemblyAI's territory; clinical-grade medical transcription or on-premises and edge deployment stay with Speechmatics.

HokAI guides covering Speechmatics

More AI Tools on HokAI

Visit Speechmatics Official Website