Pick the language before you pick the line
There is a running order most teams use for short video, and it has a bug in it.
Someone writes the English line. The clip gets generated, reviewed, approved. Then a market lead asks for the Spanish version, and the request goes back as if it were a subtitle job. It is not a subtitle job. On any model that generates voice in the same pass as the picture, a language change is a new take, not a new audio track laid over an old one.
That distinction is worth spelling out, because it changes when the decision has to be made.
Why the language is upstream of everything
When speech is generated jointly with the image, the mouth shapes, the head movement, the breath before the sentence, and the small pause the actor takes at the comma are all downstream of the phonemes. Change the phonemes and every one of those changes with it. You do not get "the same clip in Spanish." You get a different performance that happens to be framed the same way.
This has two practical consequences.
You cannot approve a clip and then localise it. What gets approved is a take. The Spanish take will differ in timing and in micro-expression, and if the approval was granular enough to matter, it has to happen again.
Your line length budget is per-language, not per-script. This is the one that actually costs people a re-render. English is compact. A twelve-word English line that lands comfortably in five seconds becomes a considerably longer utterance in Spanish, German, or Japanese, and it will either get rushed or run past the end of the clip. Neither failure looks like a language problem in review. It looks like "the pacing is off."
Count syllables, not words
The fix is boring and it works: write the line in the target language first, read it out loud at delivery pace, and time it. If you need a rule of thumb before you have the translation, count syllables. Roughly three to four syllables per second is a natural conversational pace; a fifteen-second ceiling gives you somewhere near fifty syllables of comfortable speech, and less than that if you want the character to breathe.
Then write the shortest version of the line that survives translation. Short lines localise well because they have fewer places to grow. A subordinate clause that adds two words in English adds six in German.
The supported-language list is a casting constraint
Every model that generates dialogue has a set of languages it is actually stable in, and a long tail it will attempt. The tail is where teams get hurt: the model produces something confident-sounding that a native speaker immediately flags as wrong-accented or wrong-stressed, and you find out at review rather than at brief.
Treat the supported list the way you would treat a casting availability sheet. If a market is not on it, the plan for that market is a different plan — a text-only cut, a voiceover recorded separately, or a card at the end. Deciding that in the brief costs nothing. Deciding it after two rounds of review costs a week.
If you want to see how that list is normally presented, the language list printed next to the field that uses it is a reasonable reference point — it shows the eleven languages held to a stable standard, alongside the prompt field where the line actually goes, which is the layout that makes the constraint hard to forget.
A brief that survives localisation
Four lines at the top of the brief, before anything visual:
- Markets, listed. Not "EU," but the actual languages.
- The line, in every target language, written by someone who speaks it.
- The longest of those, timed out loud, in seconds.
- The clip length, set from item 3 rather than from the storyboard.
Everything visual gets planned against the longest line, because the shortest one can always sit in silence for half a second and look deliberate. The reverse is not true — a line that overruns has nowhere to go.
The part that does not localise
One last thing worth flagging, because it surprises people: the ambience does not need to change. Room tone, footsteps, a kettle, traffic — these read the same in every market, and keeping them fixed across language versions is what makes the set of clips feel like one campaign rather than five unrelated ones. Change the voice, keep the room.
If you want to try the sequencing on something small before committing a campaign to it, generating the same eight-second clip in two languages and watching the timing drift is a five-minute exercise; a browser-based generator such as minimax-h3ai.video is enough for that, and the drift is obvious the moment you put the two takes side by side.
Language first. Line second. Frame last.
All Rights Reserved