Guide
We Run Seedance, Kling, Veo and Grok in Production. Here Is When Each One Wins.
“Cheap per second is not cheap per usable second. We price per finished film because that is the only number a client can actually spend.”
Which engine wins depends on the shot, not on the engine. As of late August 2026 the split at Filmito Studio is this: Seedance gets the shots that need convincing physics and a face that stays the same face, Grok Imagine gets the shots that start from an AI-generated still, Kling gets the shots we can frame tightly in the prompt, and Veo gets the next attempt when a shot has tripped a known blocker on one of the other three. That is the short answer. The rest is the reasoning, including where each engine has cost us a take or a day.
One caveat first. This is studio experience, not a benchmark. We have not run controlled tests, we do not keep a scoreboard, and these models change roughly every month, so this is a snapshot of how we work today. If you want a ranked table, the 2026 comparisons from Magic Hour and Atlas Cloud are the kind to read. This is the other thing: what it is like to depend on four engines when an ad is due on Thursday.
Seedance: physics, faces, and the rules it taught us
Seedance, now at 2.5, is where most of our hero shots start. Of the four engines we run, it has the strongest physics and the strongest character consistency. A hand gripping a bottle, cloth moving with a body, a face that has to be the same person across a dozen shots: those go to Seedance by default. It also has the longest list of rules, each learned by losing a take.
- It rejects AI-generated images as reference inputs; it wants real photographs or stills. If a client has no photo of the product, we cannot generate one and feed it in. That shot goes elsewhere.
- On-screen text inside generated footage garbles: labels, signage, packaging type. So we generate a clean plate with no text and add the type in post. Clean plate, then post, is a house rule now.
- Moderation can refuse a prompt, so we run a free preflight check before any charge. If a prompt will be refused, nobody pays for it.
- A single take holds reliably for roughly 5 to 8 seconds. Past that, it drifts. So we cut films from short takes, the way any edit is cut.
The job we hand it: product in hand, recurring characters, and any shot where the audience would notice a body doing something a body does not do.
Kling v3: lock the frame or lose the subject
Kling is the second serious engine in the rotation, and its value to us is partly that its failures are different from Seedance's. The shot that breaks on one often survives on the other.
Two things we know for certain about Kling v3. First, subject framing has to be locked in the prompt. Say what the subject is and where it sits in the frame and Kling holds it; leave room and it wanders off-subject into a beautiful shot of the wrong thing. Second, there is a cap of roughly five parallel generations, which sounds like a technicality until you need twenty variations before a review. It changes how we schedule a day.
The job we hand it: shots we can describe tightly, and second reads on shots that Seedance's preflight or its text problem has ruled out. We will not list strengths beyond what we have verified in our own work.
Grok Imagine 1.5: the one that accepts an AI still
Grok Imagine 1.5 does one thing that decides a lot of shots for us: it accepts AI-generated images as the input for image-to-video. Seedance will not. That single difference is most of why Grok is in the rotation.
Plenty of shots involve something that does not exist as a photograph: a scene from another century, a character the client invented, a product colour not yet manufactured. For those we build a still first, iterate while it is cheap, get the client to sign off on a frozen frame, and only then set it moving. Grok takes that approved still and animates it.
The gotcha is price. Grok sits at the same cost floor as Kling, so it is not the bargain lane.
Veo 3: the engine we say the least about, on purpose
Veo 3 is in our production rotation, and this is the shortest section, because our notes on it are thinner than for the other three and we would rather write less than pad. If a post ranks Veo first or last with confidence, ask how many client films that confidence rests on.
The job we hand it today: the next attempt when a shot has tripped a known blocker elsewhere, and a fresh read when two engines have given us two versions of the same problem. When our notes firm up, we will update this post and say so.
Why per-second price comparisons mislead
The premium engines have a cost floor of roughly $0.14 per generated second on Grok and Kling. Multiply that by a 30-second ad and the whole industry looks like a rounding error. That number is wrong, and understanding why is the most useful thing in this post.
A cinematic ad uses many takes per usable second. A prompt gets refused at preflight. A take drifts at second nine. A hand is wrong on one take; the face changes on the next. Text garbled, so you generate the clean plate again. Every one of those is a generated second you paid for and did not use. The ratio of generated seconds to finished seconds is the real cost of an AI ad, and nobody prints it on a pricing page.
For scale, AI generation compute at Filmito runs about EUR 5,000 a month. That is what four engines cost to operate at a studio delivering finished films, and why we price per finished film, not per second. The full breakdown is at filmito.io/en/ai-video-ad-cost.
“Cheap per second is not cheap per usable second. We price per finished film because that is the only number a client can actually spend.”
This is also where we are honestly not the right choice for some people. If you enjoy running these engines yourself, if your time is free, and if you can treat refused prompts and drifted takes as tuition, use the tools directly. It will be cheaper in cash. We only make sense when the finished film is what you need and the takes are a cost you would rather not carry.
How a director chooses, shot by shot
Nobody at the studio picks an engine for a film. We pick one for a shot, starting from the storyboard, not the tools. For each shot the questions are the same:
- Does a real photograph of the subject exist? If yes, Seedance is open. If the subject only exists as an AI still, Grok is the natural first engine.
- Does a character recur? If a face has to hold across shots, that pushes hard toward Seedance and toward cutting the scene from 5 to 8 second takes.
- Is there text in the frame? Then whichever engine we use, the plan is clean plate first and type in post.
- Can we describe the subject and its place in the frame precisely? If yes, Kling is a strong candidate. If the prompt must stay loose, we expect Kling to wander.
- Has this shot already failed somewhere? Then it goes to the engine with a different failure profile, and Veo gets its look.
A finished 30-second ad might carry shots from three engines, and the viewer never knows, because the edit is what they see, not the generation. Check that against the films at filmito.io/en/ai-video-ad-examples; we do not label which engine made which shot, and that is rather the point.
Everything above is true of Filmito Studio as of late August 2026. Some of it will be wrong within a few months, because the engines will have changed, and when that happens we will change this post. If you want us to direct an ad across all four engines, filmito.io/en/see-your-ad is where to send the idea. If you would rather run them yourself, we hope the rules above save you a few takes.