How AI dubbing actually works (when you refuse the magic button)
Dubbing is a pipeline with review points, not a black box that 'just translates the video.'
Last updated August 5, 2026
The magic button is the wrong mental model
Search "AI dubbing" and you will mostly get tool roundups: upload a talking-head video, pick a language, export a lip-synced cut. Some of those products are real. Many of the pages still talk as if localization were one opaque job with one quality score.
For creators shipping podcasts, course audio, explainers, and product narration, the unit that matters is the script. Bad names in the transcript become bad names in the new language. A literal translation that is too long for the scene will still be too long after a pretty mouth warp. If you cannot stop the chain and edit the words, you are hoping the model guessed right.
We built Dubbing & Translation as three live steps you can inspect: transcribe, translate, re-voice. This post is the craft version of that pipeline: what each step is for, what breaks, when subtitles win instead, and where our product honestly stops (audio out, no video lip-sync).
What "AI dubbing" means in plain terms
Dubbing replaces the spoken track with a new performance in another language. Subtitling keeps the original performance and puts translated text on screen. The trade-offs are not fashion; they are effort, cost, and format. We keep a short glossary at dubbing vs subtitling.
An AI dub is not a single model that "understands video." In production stacks it is almost always a chain:
- Speech-to-text turns the source audio into words (and, in better tools, speaker labels and timing).
- Machine translation (plus human post-edit when quality matters) turns those words into the target language.
- Text-to-speech re-voices the translated script. Optional extras then stretch timing, clone a speaker, or warp mouths in video.
Step 1: Transcribe like the next person will rewrite you
Everything downstream inherits the transcript. If the model mishears a product name, a dosage, or a proper noun, the translation will confidently launder the error into another language. That is why our first step is a Whisper-class speech-to-text pass that returns editable plain text, the same transcription layer behind Speech to Text.
House limits we actually ship: common audio containers up to a published size cap, optional browser recording with a short take limit, and plain text without word-level timestamps today. We would rather say "no timestamps yet" than invent a subtitle file with fake timing. Fix the words before you translate. Expand acronyms the ear will mangle. Mark multi-speaker turns by hand if the take is a conversation.
- Listen once while reading the transcript. Correct names first.
- Cut filler only when it does not carry tone you need in the target language.
- If the source is noisy or heavily accented English, expect more manual cleanup. English-source audio is still the most reliable starting point for this pipeline.
Step 2: Translate for the ear, not the dictionary
Machine translation is fast and incomplete. Good localization shortens lines that will not fit the breath, swaps idioms, and keeps register honest for the audience (a kids' explainer is not a legal notice). Our translate step targets a fixed set of languages live in the studio: Spanish, French, German, Italian, Portuguese, Japanese, Hindi, and Arabic. The full how-to lives in the dubbing guide.
Bracketed stage directions stay in place through translation on purpose. A cue like [softly] or [pause] is performance metadata, not prose. If you strip it, the re-voice loses the turn you meant to keep. If you leave a bad cue in, the cue-following engine will still try to act it. Review both the words and the directions before you spend a keeper render.
- Read the translation aloud once. If you stumble, the listener will too.
- Prefer shorter clauses when the target language runs longer than English.
- Keep product and legal terms consistent with your glossary, not with whatever the model invents mid-paragraph.
- Rewrite dialect-sensitive lines yourself when the market needs it (for example, Latin American Spanish vs Spain Spanish conventions).
Step 3: Re-voice on an engine that actually speaks the language
The final speech is ordinary text-to-speech on a multilingual path. On Cantari that engine is Gemini Flash: it is the roster member registered for non-English work and for bracketed cues. Kokoro, Grok Voice, MAI Voice 2, and Zonos are English-only today, so offering them for a Spanish or Japanese dub would be theater. We print that constraint on the tool page rather than hide it.
Quality and speed still matter for the keeper take. Third-party Quality Elo for Gemini Flash is 1225.13. Our own wall-clock median to full audio on the standard sample is 2,770 ms. Those figures are the same ones on the open benchmark and the engines page; they are not a dubbing endurance score, and they do not grade translation quality.
Pick a voice once per project when you can. A course that changes cast body every module feels accidental. Export MP3 or WAV you own, then drop the track into your editor the same way you would a human VO delivery.
* Quality Elo from the Artificial Analysis Speech Arena, retrieved 2026-06-10. User-vote arena ratings, not our scores.
* Latency: our own wall-clock measurement to full audio on the same routed path that serves the studio, median of 3 runs, measured 2026-06-10. Not a server SLA.
* Re-voice engine and language coverage follow the live dubbing tool and engine registry, not a marketing matrix.
What this pipeline does not do
Honesty is part of the craft. Our dubbing flow is audio-first. It does not retarget mouths, time-stretch video, or rewrite picture cuts. If you upload a video file, we read the audio track and still return audio. You bring the dubbed track back into your NLE yourself. That is the same stance as the product docs, not a temporary footnote.
We also do not ship subtitle files from this path today. Transcripts lack honest timestamps, and a .srt with invented timing is worse than no file. When you need on-screen text more than a new voice, subtitle (or a hybrid) may be the correct product choice even if "AI dubbing" is the hotter search phrase.
Multi-speaker panels, heavy music beds under dialogue, and fast overlapping talk still need human cleanup. A composite pipeline reduces handoffs; it does not fire the editor.
When to dub, when to stop
Dub when the format is audio-first, when the audience will not read along, or when a second language version has to stand alone (localized podcasts, training audio, radio-style ads). Keep the original and subtitle when performance identity is the product, when budget is thin, or when viewers commonly watch muted.
For market-facing video that already has a face-led cut, a separate lip-sync or retiming tool may sit after a clean script pipeline. Do not let the video effect hide a broken translation. Fix words first. The localization use-case story we keep next to the tool is dubbing and localization for video.
A practical sequence you can run this week
If you want a checklist instead of theory, use this order on a five to ten minute source.
- Export a clean mono or stereo dialogue stem if you can; less bed music under the voice means fewer STT errors.
- Run transcription. Fix every proper noun and number before you touch translation.
- Translate into one target language you can actually review (or have a native reviewer).
- Shorten lines that sprawl. Keep bracketed cues only where a director would mark a turn.
- Re-voice on Gemini Flash. Audition two voices max, then lock one.
- Spot-check the worst risk lines (intro, legal, product names) at 1.0x, not only at 1.5x.
- Export WAV or MP3 you own. Lay the track under picture in your editor. Stretch or cut picture only after the words are right.
- If the market needs both, add subtitles as a separate pass with real timing tools.
Honest limits
Eight translation targets are live today, not "every language on earth." Gemini Flash speaks a wider set as a TTS engine than the dubbing translation step currently offers; the tool page is the source of truth when those lists differ. English-source audio remains the most reliable intake.
Only the multilingual cue-following engine is offered for non-English re-voice on this roster. If a safety classifier refuses a lawful prompt, you still need an alternate wording, not a pretend English-only fallback for a Spanish line.
Premium generation still sits inside the flat monthly allowance (published free-tier premium budget: 10,000 characters, about 10 minutes by house math). Transcription and translation are part of the studio flow; the re-voice is the character-metered keeper step. Paid plans raise the flat allowance rather than putting you on a per-character anxiety meter.
Arena Elo and our latency sample do not measure lip-sync, cultural fit, or multi-hour course catalogs. For engine receipts, stay on benchmark and engines. For ownership of exports, see you own what you make here.
Run one real segment
The shortest proof is not another roundup list. Take a two-minute English segment you already ship. Open Dubbing, transcribe it, fix two names on purpose, translate into a language you can judge, and re-voice only after the script reads clean. Compare that to a one-click demo that never showed you the words.
If you need the standalone transcription step first, use Speech to Text. If you are writing target-language lines from scratch instead of from a recording, skip straight to Text to Speech on the multilingual engine. When you want the full studio, create a free account. Localization is review discipline plus an honest pipeline. The magic button can wait.
Quality Elo data: third-party, from the Artificial Analysis Speech Arena, retrieved 2026-06-10. Latency figures are our own measured wall-clock numbers, not a server SLA.
Check our work, then make your own.
The benchmark is live and the studio is free to start. Every claim above is one click from its source.
