Included by default
Sovereign Cloud
- Hosting by EU companies only
- AWS, Google Cloud, Microsoft Azure and Lambda are not engaged
- No route for a US CLOUD Act order to reach your audio, text or voice profiles
- Included by default for EU/EEA customers
Voice AI Made in Europe
EU Sovereign Cloud included on every plan, no enterprise contract.
“Good afternoon, your appointment is on 03/15/2025 at 2:30 p.m. The office is located at 42a Bridge Street. For questions, reach us at 02079460123 or info@clinic.co.uk.”
Direct comparison
Both companies build real-time speech. Sovereignty, German and deployment decide it.
| Decision criterion | KugelAudio | Cartesia |
|---|---|---|
| EU Sovereign Cloud | Included on every plan | Not offered |
| Entity and governing law | German GmbH, EU law | US company, CLOUD Act applies |
| Time to first audio | Kugel v3: 91 ms | Sonic 3.5: 111 ms · Sonic Turbo: 127 ms |
| German numbers, IBANs and emails | Correct by default | No German-specific rules |
| German voices | 112 voices, regional variants | From a multilingual model |
| On-premise deployment | Live within days | Enterprise tier only |
Cartesia figures come from its public documentation and our benchmark run, August 2026. Please verify current terms for your plan.
Benchmarks
Lower is better
Measured over WebSocket streaming from an EU data centre. Network round-trip included, connection setup excluded.
339 evaluations · OpenSkill ranking
| Model | Score | Win rate |
|---|---|---|
| KugelAudio | 26 | 78.0% |
| ElevenLabs Multilingual v2 | 25 | 62.2% |
| ElevenLabs v3 | 21 | 65.3% |
| Cartesia | 21 | 59.1% |
| VibeVoice | 10 | 28.8% |
| CosyVoice v3 | 9 | 14.2% |
Listeners heard a reference voice, then compared two models and picked the one that sounded more human and closer to the original. German samples covered neutral speech, shouting, singing and slurred speech.
Data sovereignty
Built, trained and hosted in the EU, outside the reach of the US CLOUD Act.
Included by default
Live within days
Custom dictionaries
Store the pronunciation once in your dictionary. From then on, the model says the name correctly in every single response.
Your text
“Order your new tableware at Villeroy & Boch.”
Your dictionary entry
Villeroy & Boch → [ˌvɪləɉɔɪ ʔʊnt ˈbɔx]
The output
Pronounced correctly. Every time, in every response.
Works the same for product names, technical terms and abbreviations, via IPA notation or a simple replacement.
Pricing
No character accounting, no seat tiers.
Fast, budget-friendly text-to-speech model
€0.035 / min
Premium text-to-speech model with maximum naturalness
€0.07 / min
On-premise, dedicated capacity and committed-usage rates are available on request.
* Inference TTFA, measured server-side. More in the latency docs · ** Character costs assume approximately 825 characters per minute.
FAQ
KugelAudio is a German company building real-time text-to-speech in 26 languages, hosted in the EU. The EU Sovereign Cloud option engages only EU companies as subprocessors, and it is included on every plan rather than sold as an enterprise add-on.
In our benchmark run, Kugel v3 reached a 91 ms median time to first audio, against 111 ms for Sonic 3.5 and 127 ms for Sonic Turbo. Measurements used WebSocket streaming from an EU data centre, with connection setup excluded.
Yes. KugelAudio is a German GmbH under EU law with its IP in Europe. Under the Sovereign Cloud agreement, AWS, Google Cloud, Microsoft Azure and Lambda are not engaged, so no provider with a US parent processes your data.
In a blind A/B test with 339 human evaluations on German samples, KugelAudio ranked first with a 78.0% win rate and Cartesia fourth with 59.1%. KugelAudio also offers 112 German voices and German normalization of numbers, IBANs and email addresses.
Yes, within days. The models run in your data centre behind your firewall with the same SDK and API as the hosted service, handling up to 192 concurrent calls per GPU.
Real-time, EU-hosted voice that scales with your product.