Skip to main content
Audio processing runs in the split step, the first stage of an AI Dubbing Order. It separates the source audio into an isolated dialogue stem and a music-and-effects (M&E) bed. That separation is what makes dubbing sound like dubbing rather than a voiceover. The dialogue stem feeds transcription; the M&E bed is preserved and remixed under the synthesized target-language voice, so the score, ambience, and sound design survive localization intact.

At a glance

It is the shortest step in the dubbing pipeline, and the one with the largest effect on whether the finished mix sounds professional.

LALAL.AI

A dedicated stem separation service built on a set of proprietary neural networks — Phoenix, Orion, Perseus, Andromeda, Lyra, and Lynx — capable of splitting a track into twelve distinct stem types. An auto option selects the most effective network for the requested stem. The Orion network takes a different approach from conventional separation. Rather than carving stems out of the mix with spectral masks, it synthesizes them, reconstructing the stem from what it has learned across many mixes. The practical result is markedly less of the artifact residue that betrays a dub: the watery, phase-smeared remainder where dialogue used to be. Best for — every AI Dubbing Order. Dialogue-over-music content benefits most, since that is where mask-based separation degrades hardest. Strengths
  • Direct-synthesis separation with lower cross-stem leakage than mask-based methods.
  • Three noise-cancelling levels on the voice stem, for source audio of varying quality.
  • Dereverberation, which removes room echo from the extracted voice. This matters beyond the mix: reverb in the source is one of the more reliable ways to degrade transcription accuracy, so cleaning it here improves every downstream step.
  • Separate voice and music stems — spoken voice is a distinct target from sung vocals, which is exactly the split a dubbing pipeline needs.
  • Selectable extraction level (deep_extraction or clear_cut), trading completeness of the extracted stem against cleanliness of what is left behind.
  • Finer stems — drums, bass, piano, guitar variants, synthesizer, strings, wind — for mixes that have to be partially rebuilt rather than simply split.
Considerations
  • Separation quality tracks source quality. Heavily compressed audio, dialogue buried under a loud mix, or material already processed by another separation tool will produce a weaker M&E bed.
  • Where the source has a clean dialogue stem available, supplying it directly will always beat separating it back out.
In OllanglalalAi · AI Dubbing · Default.

Text-to-Speech providers

The voice engines that render into the preserved M&E bed.

Source files and assets

Supported source formats and how audio is ingested.