Voices, languages, and a song if you want one

/ The short version:
- Characters get voices of their own; narration is a separate track from dialogue.
- Twelve languages are spoken rather than subtitled, and the film is written in them rather than translated at the end.
- Lip sync is switched on per scene, because it does nothing at all on a shot with nobody speaking in it.
- Sound effects are written down per scene in the same way the pictures are, which is what makes them reviewable before they exist.
The quickest way to make a film feel automated is to have one voice read every part. Characters get voices of their own here, and narration is a separate track from dialogue rather than the same track with a different label.
Sound is the stage most generative video tools reduce to a music dropdown, which is odd, because sound is more than half of what makes something feel finished. A rough cut with good sound reads as a film. A beautiful cut with one flat narrator reads as a demo.
/ In this post:
A voice per character
Because characters are first-class objects — see why the faces hold — a voice is something a character has, not something a line has. Cast a character once and every line they speak, in every scene, is spoken by them.
- A voice per character — chosen on the cast record, not per line, so it cannot drift between scenes any more than the face can.
- Narration separate from dialogue — different track, different treatment, different behaviour under the music.
- Your own voice — cloned from a recording, then used for every line that character speaks.
- Lip sync where it matters — switched on per scene, because lip sync on a shot of a landscape does nothing at all.
Twelve languages, spoken rather than subtitled
The distinction is not cosmetic. Subtitling a film means generating it in one language and putting text over it; the performance, the timing and the lip movement all belong to the original. Writing it in the language means the narration is composed in that language, spoken by a voice appropriate to it, and timed to it.
The language is chosen at the storyline stage, before anything is written, for exactly this reason — it is a decision about what the film is, not a post-process.
What we hear, written down
Every storyboard card carries a “what we hear” field alongside “what we see”. That is not decoration; it is what makes sound effects reviewable. A field you can read is a field you can correct before it is made, in the same way and for the same reason the rest of the storyboard is text.
Ducking is automatic and worth calling out because its absence is the single most common giveaway in amateur edits: a good music bed at a constant level, and a narrator fighting it for the whole minute.
The score ducks under anyone speaking, because a film where you cannot hear the dialogue is not a film.
Directing the score
Five moods, eight instruments and three tempos, or describe what you want in your own words. The point of having both is that you should not have to use words when a menu is faster, and should not be stuck with the menu when it does not cover what you mean.
Or make the whole thing a song
If the film is a song, the lyrics are written for you across fourteen genres and sung, rather than a backing track with a voice over the top of it. The lyrics are versioned like everything else, so you can go back to a verse you preferred.
And the captions come from the recording
Once the narration exists, the timing of every word in it is known. Captions are built from those timings rather than guessed from the script or divided evenly across the shot. That is a sound-stage fact with a visible consequence, and it is covered properly in why your titles read as typed.
Questions
Can each character have a different AI voice?
Yes. A voice is a property of the cast member rather than of a line, so every line a character speaks in every scene is spoken by the same voice. Narration is a separate track from dialogue.
How many languages are supported?
Twelve, and the film is written and spoken in them rather than subtitled afterward. The language is chosen at the storyline stage, before anything is written, because it changes what gets written rather than what gets displayed.
Can I clone my own voice?
Yes, from a recording of about thirty seconds. The clone is then used for every line that character speaks, in every scene.
Is lip sync applied to the whole film?
No — it is switched on per scene. Lip sync on a shot with nobody speaking in it does nothing, so it is a per-scene decision rather than a setting for the film.


