Color and product accuracy in AI video: when "close enough" isn't
This update was drafted on a schedule by the AI I build with, from real project notes — part of the vibecoding experiment this blog documents.
Let me stay on this thread for one more post, because I think I only told you half of it.
Last time I argued that text belongs in the edit, not the prompt — the model paints something letter-shaped, and letter-shaped isn't a brand name. Fine. Solved. But then I went back through the Cadence test renders with the text problem mentally crossed off, and the shots were still wrong. Not wrong like a typo. Wrong like: that is not the color of that can.
And that one's sneakier, because there's no obvious error to point at. Nobody looks at a slightly-off teal and says "typo." They just feel that something's a bit cheap. Which, if you're handing it to a client, is somehow worse.
The thing nobody warns you about
Text drift is loud. You read the label, the letters are wrong, you kill the shot. Takes a second.
Color drift is quiet. The model gives you a beautiful, confident, plausible version of your product. The blue is a blue. The bottle is a bottle. The sneaker has the right general silhouette. Everything is in the neighborhood of correct, and that neighborhood is big enough to hide in.
So you approve it. And then you put it next to the client's actual packshot and the difference is immediate and embarrassing.
This is the same thing I keep running into everywhere, honestly. It's confidence without sufficiency wearing a different outfit. The output doesn't look uncertain when it's wrong. It never does. There's no visual tell for "I guessed."
Why the model can't just get the color right
Two reasons, and they're worth separating because they need different fixes.
One: the model isn't reproducing, it's producing. A hex code is an exact number. The model doesn't take a number and fill pixels with it. It's generating something that reads as, say, deep teal, in the lighting it decided the scene has. "Deep teal" covers an enormous range of actual values, and the model will happily pick a different point in that range for each generation. Not a bug. Just what the thing is.
Two: lighting legitimately changes color, and you can't tell the two apart. Here's what makes this hard to even diagnose. In a real shot, a real can under warm key light genuinely isn't the same RGB value as that can on white. That's physics, and it's supposed to happen. So when a generated frame comes back off-color, you can't immediately tell whether the model got the color wrong, or got the color right and lit it in a way that shifts it. Both look identical in the frame. You only find out which when you try to correct it and it fights you.
Product shape has the same two failure modes, minus the excuse. The number of ridges on the bottle cap. Where the seam sits. The exact proportion of a sneaker's sole. The model is pattern-matching to "sneaker," not reproducing that sneaker, and the drift shows up in exactly the details a brand's actual customers recognize instantly.
Prompting doesn't fix it (I tried for a while)
I went through the obvious moves. Hex codes in the prompt. Pantone names. "Exact brand color." Long descriptive paragraphs about the shape of the cap. Reference images.
Reference images help the most, by a wide margin — the same way reference frames help continuity. They pull the output toward the right neighborhood. But "toward" is the operative word. It's still an influence on a generation, not a constraint on one. And a shot that's 90% of the way to the right teal is still a shot you can't ship, and you can't re-roll your way to a number.
That's the signal I set for myself when I was writing up re-roll budgets: when re-rolling is the plan, the plan is wrong. Color hit that signal fast.
So what actually works
Same shape as the text answer, which I guess is the real lesson here. Let the model do the part it's extraordinary at, and let a deterministic tool do the part that has to be exact.
Grade it instead of prompting it. Color correction is a solved problem that predates all of this. You generate the shot, then you pull the product's color to the right value in the edit. This is a normal, boring, well-understood step in every real production pipeline. Secondary correction — isolating one object's color without touching the rest of the frame — is literally what the tooling was built for.
Prompt for gradeable, not for correct. This is the part I had to learn by getting it wrong. It's much easier to grade a shot that's evenly lit and roughly right than one that's dramatically lit and badly wrong. So I stopped writing prompts that fight for the exact color and started writing prompts that produce a clean, controlled base: neutral-ish light, product clearly separated from the background, no wild colored spill bouncing off it. Boring to generate, easy to fix. A shot with a neon pink practical light smeared across the product is unfixable in the grade, no matter how cool it looks.
Composite the product when the shape has to be exact. If the client's product has a recognizable form and the model keeps approximating it, the honest answer is the model doesn't render that product at all. Generate the environment, the light, the motion — the stuff it's genuinely great at — and put the real product in. Same split as the logo. Same tracking job, too.
Treat the brand's assets as the source of truth, not the prompt. The hex value, the vector, the product photos. The prompt is a request. The asset is a fact. Anything that has to be exact should come from the fact.
The line I use
I ended the text post with a one-liner and I'm going to do the same thing here, because it turns out it generalizes:
If being a little bit off is a problem, it doesn't come from the model.
Letters. Brand colors. Product geometry. Prices. Anything with a single correct answer that someone can check. The model is for the things with a range of good answers — light, atmosphere, motion, feel — and it is genuinely, unreasonably good at those. That's the trade, and it's a good one once you stop fighting it.
Where I'm actually at with this
Being straight about it, same as last time: this comes from building and stress-testing the pipeline, not from a pile of delivered client work. The four reels on the Cadence page — AURA, AXIOM, LUMEN, STRIDE — are invented brands, each a single 8-second Seedance 2.0 clip. Nothing in them had to match a real product, and that is exactly why they could be pure generations. The test only gets real when it's someone's actual can, in their actual color, next to their actual packshot.
And look, this may be a shrinking problem. Native color and reference control keep getting better, and the day a model holds an exact hex and an exact silhouette across a multi-shot sequence, half this post evaporates. I'd be happy about it.
Until then: the model makes the world, the edit writes on it, and the grade makes it the right color.
This one's auto-drafted from my notes on a schedule. If a number isn't in the notes, it doesn't show up here — I'd rather leave a blank than make something up.