0

Localising a generated clip when the audio has no stems

You have a thirty-second spot generated by a video model, it is approved, and now somebody wants it in four more languages.

With conventionally produced footage this is a solved problem: the picture is one asset, the voice is another, the music is a third, and localisation swaps one of the three. With footage from a model that generates speech, effects and music in the same pass as the picture, that separation does not exist. There is one file and it is a mix.

This changes the architecture of a localisation workflow enough to be worth planning before the first clip is approved rather than after.

The constraint, stated precisely

Three facts, all of them properties of the model rather than of any interface on top:

  1. Audio is generated with the picture in one pass. There is no separate audio render to re-run.
  2. No stems are returned. You get the mix. You cannot attenuate the music, isolate the voice, or replace one effect.
  3. Turning audio off does not reduce the price, because there is no second pass to skip.

Together those mean: changing a spoken line means generating again. Not re-mixing, not re-dubbing over the existing file — generating.

Three strategies, and when each is right

Strategy A: generate per language

Run the same prompt once per target language, with the spoken line changed.

  • Cost: linear in languages. Five languages is five generations.
  • Risk: the picture will not match across versions. These models are not deterministic, so you get five similar but distinct films. For a product shot that is usually unacceptable; for a mood piece it is often fine.
  • Mitigation: cite the same reference images in every run. That pins the subject — a face, a product, a space — much more tightly than the prompt alone. It does not pin the camera move.

Use this when the deliverable is one shot, the languages are few, and visual identity across versions is not contractual.

Strategy B: generate silent-intent, add voice in post

Write the prompt so that nobody speaks on camera — narration-free action, no lip movement — then lay your own VO and music underneath in an editor.

  • Cost: one generation total, plus normal audio production.
  • Risk: you lose the thing that made same-pass generation attractive. The score no longer sits inside the scene; you are back to library music and alignment work.
  • Note: you cannot save money by disabling audio on the generation, so the model's mix comes back regardless. Treat it as a scratch track — it is genuinely useful as a timing reference for where the beats fell.

Use this when the same film has to ship in more than about three languages, or when the voice has to be a specific person.

Strategy C: edit the line rather than regenerate the shot

Some models expose an editing mode that rewrites part of existing footage — a subject, a background, a style, or a spoken line — without regenerating the whole take. Wan 3.0 has one.

  • Cost: one generation per edit, but you keep the take you liked.
  • Risk: the seam. What is preserved and what shifts depends on the shot, and no documentation will tell you in advance.
  • How to find out cheaply: run the edit at the lowest resolution tier first. At a quarter of the top-tier per-second rate, two probes cost less than one final render, and what you are checking is not image quality — it is whether the edit held.

Use this when the approved take is specifically the thing you cannot afford to lose.

The planning consequence

The decision has to be made before the first clip is approved, because Strategy B requires writing a different prompt. Discovering the requirement afterwards forces you into A or C, and A is the one that quietly multiplies the bill.

A workable default for a multi-language brief:

  1. Decide the language count up front and put it in the brief.
  2. If it is more than three, prompt for a silent-intent picture from the start.
  3. If it is three or fewer, prompt normally and budget one full generation per language, with a shared reference set to hold the subject.
  4. Either way, run the first version at the cheapest resolution and confirm the beats land before spending on a final.

Two adjacent numbers worth having in the estimate

Reference video is billed for its own duration, at your output rate, on top of the output seconds — and it counts against the same thirty-second ceiling. Reference images are free. So a localisation workflow built on re-feeding the approved clip as a reference is the most expensive shape available; one built on re-citing the same still images is the cheapest.

Omitting the resolution field selects the highest tier, which is four times the cheapest rate. In a batch workflow generating five language variants, that default is the difference between one bill and four.


The constraints described here are published, with a source link and a verification date on each row, at the Wan 3.0 specification page. If you want to test how well an edit holds on your own footage, wan-3.run runs the model from a browser with no cloud account and the first clip free.

wan-3.run is an independent third-party interface built on Wan 3.0. It is not affiliated with, endorsed by, or sponsored by Alibaba Group or Alibaba Cloud.


All rights reserved

Viblo
Hãy đăng ký một tài khoản Viblo để nhận được nhiều bài viết thú vị hơn.
Đăng kí