Available in Ollang Workflows as
assemblyai — Closed Captions, Subtitle Translation, and AI Dubbing. See the Speech-to-Text catalog.Models
The model line splits along a breadth-versus-accuracy axis:
Universal-3.5 Pro handles speakers who slip between languages mid-sentence — English–Spanish code-mixing is called out explicitly — without breaking the transcript into separate language segments.
Key capabilities
- 99 languages on Universal-2 — the highest published language count in this catalog. AWS Transcribe lists more entries, but many are locale variants of the same language.
- Custom vocabulary — measurably improves handling of product names, people names, and domain jargon, which is where generic transcription most often fails a localization workflow.
- Speaker diarization in the same pass, so speaker turns arrive with the transcript rather than needing a second alignment step.
- Speech intelligence — sentiment analysis, topic detection, and PII redaction for regulated or sensitive content.
- Automatic language detection for mixed or unknown-language source material.
Where it fits in a workflow
This is the default choice for volume on reasonable audio: podcasts, meetings, interviews, broadcast content, and back catalogs. On clean single-speaker material its transcripts are comparable to Agentic STT without the multi-agent pass. It sits alongside ElevenLabs Scribe v2 as the other single-model workhorse. The split between them is straightforward: AssemblyAI for mainstream languages and speech intelligence, Scribe for under-resourced languages, high speaker counts, and audio-event tagging.Trade-offs
- The highest-accuracy tier covers 21 languages, not 99. Confirm which tier serves your language before assuming top-tier accuracy.
- Accuracy drops on heavily accented speech, noisy environments, and messy conversational audio — as it does for every single-pass model. That is where Agentic STT or Speechmatics earn their premium.