Available in Ollang Workflows as
openaiTts — AI Dubbing. See the Text-to-Speech catalog.Models and voices
The 13 built-in voices are alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, and cedar.
tts-1 and tts-1-hd support the subset alloy, ash, coral, echo, fable, onyx, nova, sage, and shimmer.
Output formats: MP3 (default), Opus, AAC, FLAC, WAV, and PCM.
Steering
Ongpt-4o-mini-tts, the instructions parameter controls:
- accent
- emotional range
- intonation
- impressions
- speed of speech
- tone
- whispering
Where it fits in a workflow
This is the engine for volume with direction: internal communications, training modules, product walkthroughs, compliance content, and first-pass drafts, where delivery needs steering but a bespoke voice identity does not. A useful pattern: run the full backlog through OpenAI TTS, review, and re-run only the segments that need a distinctive voice through Cartesia or ElevenLabs.Trade-offs
- No custom or cloned voices. A character cannot be carried across a series with a distinctive voice identity the way ElevenLabs, Cartesia, or Azure custom neural voices allow. If brand voice consistency is a requirement, this is the disqualifying constraint.
- Language support broadly follows Whisper’s 50+ languages, but the voices are optimized for English. Quality outside English is less consistent than the language count suggests — benchmark your specific target languages rather than assuming parity.
- Narrower expressive ceiling than Gemini TTS or Cartesia on performed content.