KugelAudio review, pricing and limits

European TTS API with fast first-audio latency across dozens of languages, GDPR-compliant EU hosting with zero data retention, and a 78% human-preference win rate over ElevenLabs on German-language tests. Billed per minute, no card required to start.

  • ai voice assistants
  • Web

Last updated: 2026-09-26

KugelAudio is a Berlin-built text-to-speech API that processes all data on EU-only infrastructure with zero retention, supporting 39 languages with time-to-first-audio around 40-50ms on its kugel-2-turbo model. It adds on-premise Kubernetes deployment for air-gapped enterprises, voice cloning from short reference clips, and a drop-in ElevenLabs-compatible proxy for teams migrating between providers.

HokAI Editorial Rating: 3.6 / 5

  • ease of use: 7.5 / 10
  • value for money: 7.8 / 10
  • support quality: 6 / 10
  • feature completeness: 7.8 / 10

About KugelAudio

KugelAudio, founded in 2025 and backed by Y Combinator's Spring 2026 batch, is a text-to-speech API built in Berlin for European enterprises that cannot accept US data jurisdiction. It was built by Kajo Kratzenstein (former CTO at Sagemode, an HPI graduate who spent two years researching TTS) and Viktor Presber (previously co-founded full-house.io to $300K ARR in six months, with Siemens Energy and Hitachi Rail as clients).

The product ships as two API models rather than one flagship: kugel-2-turbo, priced at 0.035 euros per generated minute with roughly 40-50ms time-to-first-audio, and kugel-3, priced at 0.07 euros per minute with roughly 60-80ms time-to-first-audio and added IPA phonetic notation support. Both are trained on speech data drawn from the YODAS2 dataset with particular attention to real-world text like addresses and phone numbers that trip up most TTS systems, and both include voice cloning, streaming and word-level timestamps at no extra charge. KugelAudio also ships an ElevenLabs-compatible proxy layer, letting teams point an existing ElevenLabs integration at KugelAudio's base URL and change only the API key and voice IDs. For enterprises that cannot send data to US cloud providers under GDPR, the CLOUD Act, or FISA Section 702, an Enterprise plan adds on-premise Kubernetes deployment, a dedicated account manager and 24/7 support at custom pricing.

KugelAudio targets enterprise voice AI teams building conversational IVR systems, customer service bots, accessibility tools, and multilingual content platforms across its 39 supported languages. It integrates with the Pipecat and LiveKit voice agent frameworks, plus a built-in VAPI custom TTS endpoint and a Cognigy connector, and an open-source version (kugelaudio-open) is available on GitHub and HuggingFace for teams that want to self-host the base model. Teams weighing a shortlist through Smart Match typically compare KugelAudio against Deepgram, AssemblyAI and PlayHT inside the AI voice assistants category.

KugelAudio is not the default pick for every voice AI project. In a blind A/B test of 339 human evaluations on German-language samples, KugelAudio's synthesis won listener preference over ElevenLabs v3 and Multilingual v2 78.0% of the time, but that result is language-specific and has not been published for other locales. For English-first, US-hosted work where GDPR is not a constraint, ElevenLabs' larger voice library and years of production hardening still make it the more mature choice. Teams also lose the comfort of a flat, easy-to-forecast monthly bill: pricing runs per generated minute across both live models, and KugelAudio does not publish an exact standing free-usage allowance beyond a no-credit-card signup, which budget-conscious teams should confirm directly before committing.

Pricing

KugelAudio bills by usage rather than a flat monthly fee: kugel-2-turbo and kugel-3 are priced per generated minute (see the plans table for exact rates), both including EU hosting, voice cloning, streaming and word-level timestamps. Enterprise customers instead get a custom quote covering self-hosted infrastructure, a named support contact around the clock, and discounted rates at volume. Sign-up itself is free with no credit card required, though KugelAudio does not publish an exact free-usage quota beyond that.

Compare against Deepgram's combined speech-to-text-plus-TTS pricing or PlayHT's flat per-seat plans if a predictable bill matters more than European data residency.

Plans and pricing
TierMonthly priceWhat it includes
kugel-2-turboCustomper-minute
kugel-3Customper-minute
EnterpriseCustom

Feature Comparison by Tier

Featurekugel-2-turbokugel-3Enterprise
Price per minute0.035 EUR0.07 EURCustom
Time to first audio~40-50ms~60-80msCustom SLA
Voice cloning✓✓✓
IPA phonetic notation—✓✓
EU Sovereign Cloud✓✓✓
On-premise deployment——✓
Dedicated account manager——✓
24/7 support——✓
Volume discounts——✓

Key Features

  • Sub-50ms Time-to-First-Audio: kugel-2-turbo generates its first audio chunk in roughly 40-50ms and kugel-3 in roughly 60-80ms; KugelAudio's own headline benchmark cites an under-39ms best-case inference time, keeping real-time voice conversations free of perceptible delay.
  • 39-Language Coverage: Supports 39 languages including German, French, Italian, Polish, Dutch, Czech and Swedish plus major Asian, Middle Eastern and South Asian languages with native-quality pronunciation.
  • Zero-Retention European Infrastructure: Text, audio and voice reference data are processed transiently across EU-region infrastructure and deleted immediately after each request; KugelAudio states it does not use customer content to train or fine-tune its models.
  • On-Premise Kubernetes Deployment: Runs inside a customer's own Kubernetes cluster with no external API calls, giving data-sovereign enterprises full control of TTS processing without third-party cloud dependency.
  • Voice Cloning: Clones any voice from a few seconds of reference audio, enabling custom branded voice creation without recording full voice datasets or extensive fine-tuning.
  • ElevenLabs-Compatible API Proxy: A drop-in compatibility layer lets teams point an existing ElevenLabs SDK integration at KugelAudio's base URL, swapping only the API key and voice IDs to migrate without rewriting client code.
  • Real-World Edge Case Training: Trained on street addresses, postal codes, phone numbers, and email address formats across its language set, outperforming general TTS models on the ambiguous text patterns that break automated reading.

Pros

  • Independent blind testing on German-language samples shows a 78.0% win rate over ElevenLabs, making KugelAudio a strong European alternative for production voice AI in at least one major market.
  • European-only infrastructure with zero data retention and no US jurisdiction exposure lets regulated teams in banking, healthcare and government skip a separate GDPR compliance review.
  • A drop-in ElevenLabs-compatible proxy plus an open-source self-hostable base model give teams two low-friction ways to adopt KugelAudio without a full integration rewrite.

Cons

  • Pricing is now purely per-minute usage with no published flat free-tier allowance; teams should confirm exact free credits before budgeting.
  • Enterprise and on-premise pricing remains contact-only, adding procurement time for cost comparisons.
  • Founded in 2025 with a 4-person team; support and SLA maturity trail larger vendors like ElevenLabs or Deepgram.

Data Handling

Training-data policy
Zero Data Retention: text, audio, prompts and voice references are processed transiently and deleted immediately after each request. Customer inputs and outputs are excluded from any future model training; the base kugel model itself was trained on the YODAS2 dataset, not customer data.
Data retention
Zero retention
Compliance
GDPR

Frequently Asked Questions

How much does KugelAudio cost in 2026?

KugelAudio bills per generated minute rather than a flat monthly fee: kugel-2-turbo runs 0.035 euros/minute (about 42 euros per 1M characters) and kugel-3 runs 0.07 euros/minute (about 85 euros per 1M characters), both with EU hosting, voice cloning and word-level timestamps included. Enterprise buyers get a custom quote instead, priced around self-hosted infrastructure and volume discounts rather than metered usage.

Does KugelAudio have a free plan?

Sign-up is free and does not require a credit card, and KugelAudio's own pricing page advertises a free-tier-included badge, but the vendor does not publish an exact free-usage quota; both live models bill per minute once usage begins. Developers wanting a fully free path can instead self-host the open-source kugelaudio-open model from GitHub.

Is KugelAudio actually GDPR-compliant?

Yes: KugelAudio runs on infrastructure across Germany, Finland, France, the Netherlands, Belgium, Poland and Sweden, processes text, audio and voice reference data transiently with no retention after each request, and states it never uses customer content to train its models. It also self-declares alignment with EU AI Act Article 50's AI-content transparency rules.

KugelAudio or ElevenLabs: which should you pick?

KugelAudio publishes its own listening study, 339 German-language evaluations, showing a 78.0% preference win over ElevenLabs v3 and Multilingual v2. ElevenLabs still has the larger voice library, more mature English-language production use, and wider documented framework support; pick KugelAudio for GDPR-strict, European-language work and ElevenLabs for English-first, US-hosted projects.

How do you integrate KugelAudio into a voice agent?

Create a free API key at kugelaudio.com, then call the TTS endpoint directly with a Bearer token, or drop in the built-in Pipecat, LiveKit or VAPI connectors. Teams migrating off ElevenLabs can instead reuse their existing SDK against KugelAudio's compatibility proxy, updating credentials and voice references but nothing else in the client code. EU-only traffic should target api.eu.kugelaudio.com directly rather than relying on the SDK's eu- key prefix.

Top Alternatives

  • ElevenLabs: ElevenLabs has the largest voice library and mature US infrastructure; KugelAudio answers with GDPR compliance, zero-retention European hosting, and faster time to first audio.
  • Deepgram: Deepgram combines speech-to-text and text-to-speech in one API, while KugelAudio concentrates on native-quality European languages under EU data sovereignty.
  • AssemblyAI: AssemblyAI is transcription with speaker diarization, whereas KugelAudio is TTS synthesis with European language coverage and GDPR compliance.
  • PlayHT: PlayHT offers transparent flat per-seat pricing without European data residency; KugelAudio bills per minute instead and adds residency and on-premise options PlayHT lacks.

HokAI guides covering KugelAudio

More AI Tools on HokAI

Visit KugelAudio Official Website