In Ollang Workflows, Azure is selectable through the combined Azure + ElevenLabs route (
azureElevenlabs). See the Text-to-Speech catalog.Voice tiers
Key capabilities
- Full SSML — precise control over pronunciation, intonation, pauses, and pacing. For dubbing this is how a line gets fitted to a timing budget rather than regenerated until it happens to fit.
- Named speaking styles — cheerful, sad, angry, excited, friendly, newscast, chat, whispering, empathetic — selectable per segment, plus role variants including child voices.
- Multilingual voices that carry one speaker identity across languages, detecting emotional cues and adjusting speaking style to match sentiment.
- Custom Neural Voice and Personal Voice for brand-specific narration trained on your own data.
- Automatic language detection with SSML
langtag support on the Dragon HD Omni tier, for passages that mix languages.
Trade-offs
- Baseline neural voices are reliable rather than expressive. They are excellent at being clear and consistent, and they are not the right choice when voice performance is the deliverable.
- On the major languages that every provider covers, Cartesia and ElevenLabs sound better for less. Azure’s case rests on voice variety, SSML control, and custom voices.