textToSpeech step and serves AI Dubbing Orders, including Lip Sync and Audio Description. It takes the translated dubbing script and produces target-language audio.
This is the step audiences judge most directly. A translation error is invisible until someone reads carefully; a wrong-sounding voice is obvious in two seconds.
At a glance
Language coverage and voice count measure different things. Gemini TTS reaches the most languages; Azure offers the most individual voices and the finest delivery control; ElevenLabs and Cartesia offer voice cloning, so their catalog is effectively unbounded but their language list is shorter. Decide which axis your program is actually constrained by.
ElevenLabs Multilingual V2
The default voice engine, and still the reference point for natural-sounding synthetic speech. Eleven Multilingual v2 is ElevenLabs’ emotionally-aware model for content work, covering 29 languages with a 10,000-character request limit. Its most useful property for localization is not raw naturalness but voice consistency: a voice keeps its identity, personality, and accent across every language it speaks, which is what lets one character sound like the same person in eight dubs. Best for — voiceover, audiobook, and narrative content in widely spoken languages, where expressiveness and voice consistency carry the deliverable. Strengths- Among the most natural-sounding output available in its main languages.
- Voice cloning, so a specific character or brand voice can be carried across an entire catalog.
- Stability and similarity settings trade consistency against expressiveness per project.
- Pronunciation dictionaries for brand and technical terms.
- 29 languages is the second-narrowest coverage in the catalog, ahead of only Deepgram. Cartesia covers 44 and Gemini TTS reaches 99.
- Quality falls off outside the main languages, where the model loses subtlety and can drift in tone within a render.
elevenlabs · AI Dubbing · Default.
ElevenLabs’ current model line also includes Eleven v3, documented at 70+ languages and positioned for emotional dialogue and multi-speaker audiobook work, and Eleven Flash v2.5 at 32 languages with roughly 75 ms latency. Ollang’s
elevenlabs provider runs Multilingual v2.Cartesia Sonic 3.6
The current front-runner on independent voice evaluation. Sonic-3.6 covers 44 languages — Odia and Urdu were added in this release, and Hinglish is explicitly supported — with sub-90 ms stated latency. Cartesia reports it as number one on Artificial Analysis’s Speech Arena leaderboard. Model IDs aresonic-3.6 for the continuously updated version and sonic-3.6-2026-08-27 for a pinned snapshot.
Best for — conversational and dialogue-heavy content, high-volume dubbing where turnaround matters, and language sets that reach past the ElevenLabs 29.
Strengths
- 15 more languages than ElevenLabs, including South Asian languages that most premium engines do not serve.
- Interprets emotional subtext from the transcript and calibrates delivery automatically, so an undirected line still lands with appropriate affect.
- Non-verbal expressions — laughter, sighs, disfluencies — written directly into the transcript rather than requiring separate markup.
- Contextual pacing, correct alphanumeric read-out, and heteronym handling: the three places synthesized narration most often gives itself away.
- Instant voice cloning from roughly 10 seconds of reference audio.
- Newer model line, so there is less accumulated production history behind it than ElevenLabs at studio-narration quality.
- Published latency is model latency, not end-to-end round trip.
cartesia · AI Dubbing.
The in-product description for this provider still references an earlier Sonic generation with 15 languages. Sonic 3.6 supports 44.
Gemini TTS
The broadest language reach in the catalog, and the strongest option for scripted multi-speaker content. Google’s Gemini TTS models cover 99 languages with 30 voices on the Gemini API, and support multi-speaker synthesis — up to two distinct speakers rendered into a single generation, each holding its own tone, pitch, and style across the exchange. Delivery is directed with natural-language prompts covering style, accent, pace, tone, and emotion. Three model versions are documented, all in preview:
Best for — rich narrative content, dramatized or two-character scenes, and any program whose language list runs past what Cartesia or ElevenLabs cover.
Strengths
- 99 languages — the widest coverage of any engine in this catalog.
- Genuine multi-speaker output in one generation, with per-character voice consistency.
- Prompted delivery: direction is written in natural language rather than encoded in parameters.
- Code-switching within sentences with appropriate intonation.
- All TTS model versions are in preview; behavior can change between releases.
- A TTS session carries a 32k-token context limit, and multi-speaker generation is capped at two speakers.
- Expressiveness depends on prompting — undirected output sits closer to the field than the ceiling suggests.
geminiElevenlabs · AI Dubbing.
Azure + ElevenLabs
A combined route that pairs Microsoft Azure neural text-to-speech with ElevenLabs in a single provider option. Azure’s contribution is catalog depth: over 200 neural voices across 90+ locales, spread across several quality tiers, with the finest delivery control of any engine here.
Best for — programs that need a specific locale, a specific age or gender profile, or a named speaking style that the single-vendor engines do not offer; and dubbing where lines must be fitted to a fixed timing budget.
Strengths
- The largest individual voice catalog here, with named speaking styles (newscast, customer service, chat, empathetic, whispering) and role variants including child voices.
- Full SSML control over pronunciation, intonation, pauses, and pacing. For dubbing this is the practical mechanism for fitting a line to picture rather than re-recording it.
- Custom Neural Voice and Personal Voice for a consistent brand or character voice.
- Enterprise compliance posture inherited from Azure.
- Baseline neural voices are reliable rather than expressive. On the major languages every provider covers, Cartesia and ElevenLabs sound better.
azureElevenlabs · AI Dubbing.
Deepgram TTS
Deepgram’s Aura-2 line, built for low-latency enterprise voice rather than performance. Aura-2 offers 88 voices across 7 languages, and its distribution is unusual: it is deep rather than wide.
English covers American, British, Australian, Irish, and Filipino accents; Spanish covers Mexican, Peninsular, Colombian, and Latin American. Voices are addressed as
[model]-[voice]-[language], for example aura-2-thalia-en.
Best for — dubbing confined to these seven languages, especially English and Spanish programs where regional accent needs to be cast precisely.
Strengths
- 38 English and 17 Spanish voices — more choice inside those two languages than any other engine here offers.
- Speaking styles spanning warm conversational, confident professional, and characterful storytelling.
- Low latency at scale, which translates into fast batch throughput for dubbing.
- Seven languages is the narrowest coverage in the catalog. Confirm your language pair before selecting it — this is the most common reason a Deepgram Workflow has to be reconfigured. French in particular has only two voices.
- Tuned for voice agents and IVR; accurate rather than expressive on performed content.
deepgram · AI Dubbing.
OpenAI TTS
The most directly steerable engine in the catalog.gpt-4o-mini-tts is steerable: alongside the text you supply an instruction describing how it should be delivered, and the model adjusts accent, emotional range, intonation, impressions, speed of speech, tone, and whispering accordingly. It ships 13 voices — alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, and cedar — and outputs MP3, Opus, AAC, FLAC, WAV, or PCM.
Best for — high-volume dubbing workloads, internal and training content, and first-pass drafts where delivery needs directing but a bespoke voice does not.
Strengths
- Instruction-based steering gives real delivery control without per-voice tuning; for dubbing it is the practical substitute for line-by-line direction.
- Low latency, which keeps large dubbing batches moving.
- Broad format support for downstream mixing.
- Preset voices only. No cloning and no custom voice identity, so a character cannot be carried across a series the way ElevenLabs, Cartesia, or Azure custom voices allow.
- Language support broadly follows Whisper’s 50+ languages, but the voices are optimized for English — quality outside English is less consistent than the language count suggests.
- Narrower expressive ceiling than Gemini TTS or Cartesia on performed content.
openaiTts · AI Dubbing.
Related
Provider catalog overview
Pipeline steps, provider tags, and status definitions.
Order types and workflows
How AI Dubbing, Lip Sync, and Audio Description Orders are processed.