Text in AI video: why logos and labels belong in the edit, not the prompt
This update was drafted on a schedule by the AI I build with, from real project notes — part of the vibecoding experiment this blog documents.
So if you've made even one AI video ad, you've probably had this moment. The shot is gorgeous. The light's right, the product turns just the way you pictured, the camera move is better than anything you'd have storyboarded. And then you look at the label. And the label says something that is almost your brand name. Close enough to recognize. Wrong enough that you can't ship it.
I hit that early working on Cadence, and I kept trying to fix it in the prompt. Spell it out. Put it in quotes. Say "exact text." Re-roll. It took me longer than I'd like to admit to land on the actual answer, which is boring: stop asking the video model to write words. Generate the shot clean, and put the text on in the edit.
That was one line in my storyboarding post — mark every element as "generated" or "applied," and anything with letters goes on the applied list. This is the long version. Why it's true, and what the edit actually looks like.
Why do AI video models get text wrong?
Because the model isn't writing. It's painting something that looks like the description. A letter, to the model, is just a shape that tends to show up in a certain kind of picture. It doesn't know your logo is a fixed sequence of glyphs that has exactly one right answer.
And text is the worst case for that kind of guessing. It's tiny, it's high-detail, and it's unforgiving. A leaf can come out a little different and nobody cares. A letter that comes out a little different is a typo. With your brand on it.
To be fair, the models have gotten better. Short, big, simple words come out right a lot more often than they used to. But "right a lot more often" isn't the bar. The bar for a logo is identical, every frame, every shot, every time. And video makes it harder than an image does, because the model has to keep that text correct while the thing it's printed on moves, turns, catches light, goes out of focus. Every frame is another chance for the M to grow an extra leg.
That's the same wall I wrote about in continuity is the wall. Text is continuity at its most unforgiving: fine detail that has to persist exactly. So it breaks first.
Why can't you just re-roll until it's right?
You can. Sometimes you'll even get a clean one. But think about what you're doing there. You're spending generations, which cost money and time, hoping the dice land on the one outcome out of many that has your exact spelling. And even when it lands, you have to check every frame, not just the thumbnail. A label that reads right at second one and drifts at second six is still broken.
Then you need a second shot. And the label has to match the first shot. Now you're rolling for two coincidences.
I put per-shot re-roll budgets on the board for exactly this reason. And text is the thing most likely to eat a budget and still come back wrong. That's the signal. When re-rolling is the plan, the plan's wrong.
What does "put it in the edit" actually mean?
It means the model does the part it's great at — shape, material, light, motion — and a normal editing tool does the part that has to be exact. Here's how I break it down.
Generate a clean plate. Prompt for the product with a blank label, a plain surface, an unmarked box. You're not hiding the brand, you're leaving a canvas for it. Models are much better at "a matte black can with a blank white band" than at your logo, because blank has no spelling to get wrong.
Leave room on purpose. If an end card needs a headline, prompt for negative space where the headline goes. Clean sky, an empty wall, a shallow-focus background. Don't make the edit fight a busy frame for somewhere to put the words.
Pick the right kind of overlay for the shot. This is where it gets practical:
- Flat, static text — end cards, headlines, prices, the call to action. Just a title layer. This is most of the text in any short ad, and it's the easy case.
- Text that sits on a moving product — the label on a can that rotates. That's a tracking job. You track the surface in the edit and pin the logo to it, so it moves and skews with the object. Planar tracking and corner pins exist precisely for this, and they're standard in normal video tools.
- Text that has to look physically lit — a logo embossed on leather, a sign in a scene. That's the hard one. You composite, then match blur, grain and light so it doesn't look stuck on. Worth it for a hero shot. Not worth it for every shot.
Keep the real artwork the source of truth. The logo that goes on is the actual vector file from the actual brand. Which means it's identical in shot one and shot twelve by definition. You don't have to hope for consistency. You get it for free.
What else do you get by keeping text out of the prompt?
This is the part I didn't expect. Separating text from the generation isn't only a fix for bad spelling. It changes what the finished ad is.
It's editable. The price changes. The promo ends. The client wants "Shop now" instead of "Order today." If the words are baked into generated pixels, every one of those is a new generation and a new round of checking. If they're a layer, it's a thirty-second change.
It's reusable. One clean plate can carry five different headlines for testing, or the same spot in three languages. Try getting five matching generations that differ only in the words.
It respects the platform. On vertical video, the app's own interface sits on top of your ad — captions, buttons, the username, the description. If text is a layer, you move it into the safe zone. If it's generated into the scene, it's wherever the model put it.
It's honest. A real brand's name and logo, rendered exactly, is the brand's own artwork. A model's approximation of it is a knockoff with the brand's name on the invoice. I'd rather deliver the first one.
When is generated text actually fine?
I'm not a purist about it. There are places where letting the model draw text is fine:
- Background texture nobody reads. A street sign blurred in the back of a shot. Newspaper in a prop pile. If no viewer is going to read it, it doesn't need to be right. Just check it doesn't accidentally say something weird.
- Type as a visual effect. Abstract letterforms flying through a kinetic tech spot, where the point is the motion, not the meaning.
- Fictional brands. The four demo reels on the Cadence page are for made-up brands, and they're single 8-second generations. Nobody's real label had to be exact, which is part of why they could be pure generations. The minute it's a paying client's actual product, that changes.
The test I use is simple: if a wrong letter would be a problem, the letters don't come from the model.
The honest caveat
I'll be straight, the same way I was in the storyboarding post. This is the method I've settled on from building and testing the pipeline, not from a stack of delivered client work. Nobody's paid me for a multi-shot spot yet. And the day a model can hold exact text across a moving, relit, multi-shot sequence, a lot of this compositing goes away. That day is coming for short words. It's further off for a real logo on a real product in every frame.
Until then, it's a clean split. The model makes the world. The edit writes on it.
This one's auto-drafted from my notes on a schedule. If a number isn't in the notes, it doesn't show up here — I'd rather leave a blank than make something up.