What one prompt-to-video pipeline can and can't do yet
This update was drafted on a schedule by the AI I build with, from real project notes — part of the vibecoding experiment this blog documents.
Let me try to draw a line that almost nobody draws honestly, because both sides of this have a reason to lie about it. The AI video people need it to be further along than it is. The people whose job it threatens need it to be worse than it is. Neither is right, and the actual line is specific and pretty easy to describe once you've spent real time on the far side of it.
I built Cadence Studios on top of a prompt-to-video pipeline — Claude plans the creative, Higgsfield renders it with Seedance 2.0. Four reels came out of it as a range test: AURA, a beverage product spot. AXIOM, a hyper-kinetic tech launch. LUMEN, a warm cinematic education promo. STRIDE, a high-energy sneaker spot. Each one eight seconds, 9:16, generated end to end. They're on the project page if you want to judge them yourself rather than take my word for it.
Here's what I actually learned.
What it does genuinely well
A single shot, and the shot looks real. This is the part that still gets me. The lighting behaves. Materials read as materials — glass looks like glass, liquid moves like liquid, fabric has weight. Camera moves have actual physical logic to them instead of that floaty drift that gave the earlier generation away instantly. Put one of these next to a stock library clip and the AI one often wins, which is a sentence I could not have written a year ago.
Range across tone. This one surprised me more than the fidelity. AURA and AXIOM and LUMEN and STRIDE aren't four flavors of the same look. Calm and product-forward, then aggressive and kinetic, then warm and human, then loud and physical. The pipeline holds a mood if you describe it precisely, and mood was the thing I most expected it to flatten.
Speed, which changes what you're willing to try. The real gift isn't that a shot is cheap. It's that a wrong shot is cheap. In normal production, an idea you're 30% sure about doesn't get made, because being wrong costs a day. Here, being wrong costs a few minutes, so the 30% ideas actually get tried — and some of them are the good ones. The economics change what enters the funnel, not just what comes out of it.
What it flatly can't do yet
Continuity across shots. This is the big one and it's not close. Generate two clips of the same product and you get two products. Similar, plausibly siblings, absolutely not identical — the label shifts, the proportions move, a highlight lands somewhere else. For a single hero shot that's irrelevant. For a thirty-second spot that cuts between three angles of the same can, it's fatal. Everything downstream of that limitation is why this is a shot generator and not an ad generator.
Text and logos. Do not ask it for your brand name. You'll get something that has the shape of typography and the confidence of a real logo and is subtly, unfixably wrong. Text goes on in the edit. That's not a workaround, that's just the correct pipeline.
Precise timing. You can't say "hold the reveal 200 milliseconds longer" or "land the cut on that beat." The model doesn't take direction in units. It takes direction in vibes, and vibes don't sync to a music track. Anything rhythmic is your job afterward.
Revision by note. This is the one that changes how you have to work, and it took me a while to accept. In a normal creative loop you say "same thing, but slower and warmer," and you get the same thing, slower and warmer. Here you re-roll and get a different thing that is slower and warmer. There's no thread of identity between generation one and generation two. You are not iterating on a shot. You are sampling from a space and hoping to land near where you were.
Which means the loop isn't refine-refine-refine. It's generate a bunch, then choose. And choosing turns out to be the entire job.
The real shape of it: it makes shots, you make the film
Once you internalize the continuity limit, the whole thing reorganizes and stops being disappointing.
You don't prompt an ad. You prompt shots, plural, generated independently, and then you assemble — sequence, trim, add the type, add the sound, cut to the beat. The model is a very fast, very good, completely amnesiac cinematographer who has never seen the rest of your project and never will. Everything the model can't do is exactly the list of things an editor does. That's not a coincidence, and I don't think it's a temporary state of affairs either.
Eight seconds sounds like a constraint until you notice it's a format. A vertical eight-second clip is the unit of short-form. It's a reel, a pre-roll, a hook, a product beauty shot. The pipeline isn't bad at long-form video so much as it's aimed at a different object entirely, and the object it's aimed at happens to be the one most businesses actually need this year.
And the selection problem is real work, not a formality. When you generate four and keep one, the value you added was the judgment. Which is the same thing I keep finding everywhere else — the AI does the making and I do the checking, and the checking is where the quality actually comes from. Same story with the client sites Cadence builds out of a business's own Google reviews: the machine does the enormous tedious middle, a person owns the last inch.
So who is this useful for right now
If you need a single striking vertical clip — a product beauty shot, a launch teaser, a hook to open a reel — this is genuinely there. Not "good for AI." Good.
If you need a coherent multi-shot commercial with a consistent product and locked timing, you need an editor, and you should use this to feed them rather than replace them. That's how Cadence actually runs: the pipeline supplies the raw shots at a volume no shoot could match, and the studio part — the sequencing, the taste, the decision about which four of these ideas deserve to exist — is still mine. That was true when I renamed it from MotionForge, and honestly the rename was me admitting exactly this. "Forge" implied a box you feed prompts to. It isn't one.
I don't have client outcomes to report. Four reels rendered, the studio's live, and nobody has paid me for motion work yet — so treat all of this as an honest report from the build, not a case study. The capability is real and the gap is real, and most of what's written about this collapses one into the other.
The gap moves, obviously. Continuity is the one to watch — the day a model can hold the same product across a cut, a lot of this post is wrong and the shape of the work changes again. Until then: it makes shots. You make the film.