Skip to main content
Azure Neural Voices is Microsoft’s text-to-speech service, and the largest voice catalog available to Ollang Workflows: over 200 neural voices across 90+ language locales, organized into several quality and capability tiers.
In Ollang Workflows, Azure is selectable through the combined Azure + ElevenLabs route (azureElevenlabs). See the Text-to-Speech catalog.

Voice tiers

Key capabilities

  • Full SSML — precise control over pronunciation, intonation, pauses, and pacing. For dubbing this is how a line gets fitted to a timing budget rather than regenerated until it happens to fit.
  • Named speaking styles — cheerful, sad, angry, excited, friendly, newscast, chat, whispering, empathetic — selectable per segment, plus role variants including child voices.
  • Multilingual voices that carry one speaker identity across languages, detecting emotional cues and adjusting speaking style to match sentiment.
  • Custom Neural Voice and Personal Voice for brand-specific narration trained on your own data.
  • Automatic language detection with SSML lang tag support on the Dragon HD Omni tier, for passages that mix languages.

Trade-offs

  • Baseline neural voices are reliable rather than expressive. They are excellent at being clear and consistent, and they are not the right choice when voice performance is the deliverable.
  • On the major languages that every provider covers, Cartesia and ElevenLabs sound better for less. Azure’s case rests on voice variety, SSML control, and custom voices.

Reference