Skip to main content
Whisper is not currently selectable in Ollang Workflows. It remains documented because it is the open-source reference baseline that most comparisons in this catalog are measured against, and because self-hosted Whisper deployments are common in customer-side pipelines. For providers you can select today, see the Speech-to-Text catalog.
Whisper is OpenAI’s open-source automatic speech recognition model, trained on 680,000 hours of multilingual and multitask supervised audio. Its breadth of training data is why it handles noisy environments, varied accents, and poor-quality recordings better than most openly available alternatives, and why it became the default self-hosted choice.

Model and coverage

Key capabilities

  • Robust across acoustic conditions — background noise, varied recording quality, and accent diversity.
  • Automatic language detection without a specified language code.
  • Self-hostable, which is the deciding factor for data-residency-constrained workloads.
  • Open ecosystem — WhisperX and Whisper-JAX add alignment, diarization, and throughput improvements on top of the base model.

Trade-offs

  • Hallucination on silence and low-quality segments is the documented failure mode. Whisper will confidently emit fluent text for audio containing none, and the output does not look like an error in the transcript. This is the primary reason hosted providers are preferred for deliverable-grade output.
  • No managed diarization, speech intelligence, or SLA in the base model — those come from the surrounding tooling you build or buy.

Reference