What survives when a performance is translated, not regenerated?
I wrote a fourteen second scene in English, generated it as AI video, then dubbed it into Hindi, Marathi, Kannada and Telugu and ran lip sync on all four.
The voice survived. The tiredness, the restraint, the accent, all of it came through in four languages I did not perform in. That is the headline, and it is the reason to dub rather than generate each language separately from text.
Every dubbed version came back the same length as the original, which means any language version drops into the same slot on a timeline without moving anything.
There are two ways to put a different language into an existing video.
You can dub, which means the tool listens to the original performance and produces the same performance in a new language, trying to keep the voice and the emotion intact. The mouth on screen still moves in the old language.
Or you can generate the speech fresh in each language from text, which gives you more control but throws the original performance away entirely.
I used dubbing, and then added a separate lip sync step on top so the mouth matches the new language.
The question was not whether it sounds good. It was what changes that you would not notice by watching.
A man in his late thirties sits in a therapy room, talking to a therapist you never see. He is clearly wealthy. He speaks five short lines:
Thirty one words. The whole scene works on understatement. He contradicts himself between lines two and three, and the last line undoes everything he has just claimed, without raising his voice once.
The important detail in that chart is the split at the bottom. Sync 3 was fed the dubbed audio and the original English video, not the dubbed video. That is what keeps the length intact and stops the tool re-timing anything. Feeding it the dubbed video instead is the obvious move and it is the wrong one.
Every language needs more time than English to say the same thing. That is not a flaw, it is just how languages work.
What I did not expect was that the clip length would not move at all. The dubbing fits the new speech into the original duration rather than letting it run long. So an editor gets to treat all five versions as interchangeable.
The cost is paid in the gaps. The extra speaking time comes out of the silence between lines, almost exactly one for one. Nothing is rushed and no words are cut. The pauses are simply shorter.
I compared the English video against the dubbed video frame by frame. They are identical. Dubbing swaps the audio and leaves the image completely alone, which means you can dub late in a project with no risk to the picture.
Sync 3 has to redraw part of the image, so I wanted to know how much of it.
Comparing the English video against a lip synced version is misleading, because the file has been re-compressed and compression noise sits everywhere in the frame. So I compared the four lip synced versions against each other instead. They came from the same source video and the same tool, so the only real difference between them is the audio driving the mouth.
The answer is that almost all of the change sits in a small box on his mouth and the beard around it. The room, the plant, the chair, his hands, his eyes are all left alone. Sync 3 is doing a very local edit and not touching the performance around it.
This one needed proving rather than assuming. A lip sync tool could easily produce generic mouth movement that looks busy without matching the actual words, and you would probably not catch it by eye in a language you do not speak.
So I tracked how much the mouth moves in each frame and compared it against the pattern of the audio.
The complication is that all four versions share the same head movement, the same jaw, the same breathing, because they came from the same source video. That shared movement drowns out what you are trying to see. So I removed it, by subtracting everything the four versions have in common and looking only at what is unique to each one.
Every language matched its own audio better than it matched any other language's audio, with no exceptions. That is the proof.
Dub, do not regenerate. The performance survives, and that is not a small thing. Generating each language separately from text to speech would have handed me four different emotions for the same dialogue.