Why your titles never come out as gibberish

/ The short version:
- Garbled lettering is not a quality problem that better models fix; it is what happens when you ask an image model to draw writing.
- Titles, captions and logos are composited at assembly instead, so they read as typed because they were typed.
- Captions are built from when each word was actually spoken, not divided evenly or guessed from the script.
- Assembly composites frames that already exist, so re-cutting, re-cropping and re-downloading are unlimited.
Everyone who has used a video model has seen text come out as something that is almost letters. It is not a bug that a better prompt fixes. Image models draw the shape of writing, and the shape of writing is not writing.
The distinction matters because it tells you whether to wait for the problem to be solved or to design around it. Generative image models have improved at rendering text and will keep improving. They will also keep being a probabilistic process applied to a task with exactly one correct output, which is a bad match no matter how good the model gets. Your company name has one spelling.
They read as typed because they were typed.
/ In this post:
So nothing that has to be read is generated
Titles, captions and your logo are composited over the film at assembly, the way an editor adds them in any conventional edit. The film is generated; the text is placed. Two different operations, done by two different things, each good at its own job.
This is the same reasoning that gives removals their own channel on the storyboard rather than a “no” in the prompt — see what a storyboard card holds. Both are cases of refusing to put a job in the prompt that the prompt is bad at.
Captions on the actual words
There are three ways to time a caption, and only one of them is right.
- Divide the line evenly across the shot — the crudest and the worst. Captions drift out of sync within a sentence and are visibly wrong by the end of a long one.
- Estimate from the script — count the syllables, guess a rate. Better, and still wrong the moment the delivery is not exactly the average — a pause, an emphasis, a long name.
- Use the recording — the narration was synthesised, so the timing of every word in it is known exactly. Build the captions from that. This is what happens here.
The third option is not more sophisticated than the others so much as it is the only one that uses information already sitting there. The timings exist as a by-product of the sound stage; throwing them away and estimating would be work.
And the caption style is drawn from the film
Rather than a stock white box. The palette of the film you actually made is used, so the captions look like they belong to the thing they are sitting on instead of like a burned-in subtitle track from a different decade.
What comes out at the end
- 4K MP4, yours to keep — never watermarked, and never a preview you have to upgrade to remove.
- Vertical and square re-crops — made from the same cut. One film, three shapes, nothing re-filmed.
- A poster painted with your cast — used for publishing, and it seeds the intro and outro so they are made of the same thing the film is.
- Intro and outro bumpers — animated from that poster, with your greeting and your lengths.
- Every edit is a version — the storyboard, the lyrics and the film all keep their history, and any earlier cut can be brought back.
Assembly composites frames that already exist, lays text over them and mixes audio. Nothing in it is generated, which is why it can be repeated without limit.
Why you can re-cut as often as you like
Filming produces new frames. Assembly does not: it takes frames that already exist, composites text over them and mixes the audio under the dialogue. Nothing in that step is generative.
Which is why re-cutting, re-cropping and re-downloading are unlimited, and why the final cut assistant is allowed to re-assemble on its own rather than having to ask first.
Questions
Why does AI-generated video produce garbled text?
Because an image model draws the shape of writing rather than writing. It is a probabilistic process applied to a task with exactly one correct output, which is a poor match regardless of how good the model becomes. The fix is architectural: do not ask it to draw text at all.
How are captions timed?
From the recording. The narration is synthesised, so the timing of every word in it is known exactly, and captions are built from those timings rather than divided evenly across the shot or estimated from the script.
Will my logo be redrawn or distorted?
No. Your logo is composited over the finished film at assembly as an image, exactly as an editor would add it. It is never passed to a model to be drawn.
Do vertical and square versions have to be generated again?
No re-filming is involved. They are re-crops of the same finished cut, made by the same assembly step, which composites frames that already exist rather than generating new ones.


