← Blog

Seedance 2.5 vs MiniMax H3: Three Costs Every Comparison Misses

Every comparison of these two engines ranks the same three things: image quality, prompt adherence, motion. Useful if you are picking a model to admire. Less useful if you are producing ad creative on a budget, because the number that decides your month is not which render looks better. It is what a finished, usable variant costs once the retries, the extra ratios and the length you actually need are all in the bill.

We render on both engines in production, which means we see invoices rather than rate cards. Three costs show up there that the published comparisons do not mention. Two of them are counter-intuitive enough that we got them wrong at first ourselves.

The rates are the easy part

Both engines publish a per-second price and both are honest about it. H3 lists 768P and 2K output and returns up to fifteen seconds per call. Seedance 2.5 is available at 480p and 720p through the API and returns up to thirty seconds as a single continuous take.

Two corrections worth making before anything else, because both circulate widely. Seedance 2.5 is not a native 4K model: that specification belongs to Seedance 2.0, announced minutes apart at the same keynote, and several comparison sites have merged the two. And H3 has no 1080p tier at all: the ladder is 768P, then straight to 2K. If you are used to specifying 1080p, neither engine has a rung where you expect one.

Take the rates as read and the two look like a simple trade: cheaper and shorter, or pricier and longer. Then you render something real.

Cost one: the video you hand in is metered too

This is the point on which the published guides openly disagree. One widely-cited pricing article states that Seedance 2.5 bills on the generated video and not on what you send in. Two others state the opposite. Both cannot be right, and if you are budgeting a campaign the difference is not academic.

Our invoices settle it. Input video duration is inside the meter. The official token formula is (input duration + output duration) x width x height x fps / 1024, and the input term is multiplied by the output's dimensions, not its own. Hand a thirty-second reference clip to a ten-second render and you pay for forty seconds of pixels at your output resolution.

What makes this easy to get wrong is that the per-token rate genuinely drops when a video reference is present, roughly to 0.6 of the no-video rate. That looks like a discount, and for short references it is one. The break-even is arithmetic: a reference video only saves you money if it is shorter than about two thirds of your output length. Above that, the extra tokens overwhelm the cheaper tier and the "discount" render costs more than a text-only one.

H3 handles the same situation more transparently and no more cheaply: a reference video is billed by its own duration at the output resolution's rate, on top of the output. Reference audio is free on both, and reference images are free up to a small count, which is the whole reason a well-built pipeline continues a shot from a still frame rather than from a clip.

Cost two: continuation segments cost more, not less

If your video is longer than a single render call, you produce it as several calls and join them. The natural assumption (ours included, in writing, until we checked) is that continuation segments are the cheap ones, because they carry a video reference and video references bill at the lower rate.

Measured against our own traffic, at matched resolution inside the same task, that is backwards. A continuation segment consumes about 2.28 times the tokens of a first segment. Multiply by the cheaper rate and it still nets out to roughly 1.39 times the cost per second of the segment that started the video.

The practical consequence is not "avoid long videos." It is that the cost of length is not linear in the way a rate card implies, and any budget model that treats segment two as a discount will understate a thirty-second video. It also reframes what a thirty-second single take is worth: not just fewer visible joins, but one segment's economics instead of two-and-a-bit.

We wrote up the seam problem itself, and why joining calls is a different engineering problem than generating one clip, in our production notes on Seedance 2.0's limits.

Cost three: nobody prices the recut

Search for either engine's pricing and you will find per-second rates, token formulas, worked examples for 5, 10, 15 and 30 second clips. You will not find anyone pricing the operation that ad testing actually depends on: taking a finished vertical video and recutting it to 3:4, 1:1 or 4:5 for other placements.

It has a price, and it follows directly from cost one. A recut replays your finished footage through a model, so on an engine that meters what it is given, an extra ratio bills above that engine's normal per-second rate. On Riffkit that shows up as absolute rates: an extra ratio is 120 credits a second on Seedance 2.0 at 720p against 100 for the first, 180 on Seedance 2.5 at 720p against 150, and 80 on MiniMax H3 at 768P against 40. The cheapest engine to render on is therefore not automatically the cheapest engine to recut on, and the gap between the two narrows.

Every extra ratio is a render of its own, so count all of them in the budget before you start.

What the announcement says, and what the API gives you

The official Seedance 2.5 announcement, dated 31 July 2026, is short and specific. Up to 30 seconds per generation, with multi-round extensions. Up to 30 images, 10 video clips and 10 audio clips as references in one pass, addressed inside the prompt as @Image 2, @Video 1. It names its own weak spots too: the physical plausibility of complex motions, and the stability of scenes where several subjects interact. It states no output resolution and no language count.

Set that beside what the production API actually gives us:

Announcement Our renders
30s per generation 30s per call, against 15s on Seedance 2.0
Resolution not stated 480p and 720p are the only tiers we can select
30 images / 10 videos / 10 audio refs we send far fewer per segment (a few images, at most one reference video and one audio clip); the ceiling has never been the constraint
Silent on aspect ratio every extra ratio is a render of its own, so count all of them in the budget before you start

Where a write-up quotes 4K output or a precise lip-sync language count, that figure came from somewhere other than the announcement, and it is not what the endpoint offers. Read the announcement and the API, in that order.

Where the premium actually earns it

Two places, both concrete.

Past fifteen seconds. A thirty-second story that comes out of one call has no boundary to hide. Below that length the premium is buying you nothing you can see, because a fifteen-second render already has no joins on either engine.

When you are juggling reference material. The two engines' reference budgets differ by roughly four times: around fifty assets on the premium side against a dozen on the cheaper one. If you are holding a character, a product, a packaging shot and a voice steady across a scene, that ceiling is a real constraint rather than a spec-sheet line.

Everything else we looked at collapsed on inspection. Native synchronised audio is not a differentiator: both engines generate speech and sound with the video, both in eleven languages, and the language sets differ only at the edges. Independent reviewers decline to name a quality winner without testing identical prompts, and the failure modes they do document point in opposite directions: the cheaper engine shows motion blur and structural drift on fast action, while the premium one has been observed morphing characters in fast sequences and resisting rapid-cut prompting. Neither reads as a general verdict.

One caveat we will state plainly rather than bury: our premium-engine mileage is internal testing so far, not delivered customer volume. The cheaper engine has real customer traffic behind it. Treat our reading of the premium tier as informed rather than proven.

The workflow this points to

If the two engines were far apart on capability, you would pick one. They are not, so the useful conclusion is not a winner. It is a division of labour:

Explore on the cheap engine. Take the winning direction to the stronger one.

Generate a dozen directions (different openings, different characters, different languages) at the rate that lets you not care whether each one lands. Most will not. That is the point of exploring, and it only works when a discarded render is cheap enough not to hurt. Then take the direction that earned it, meaning the source, the creative brief and the product, and render it on the engine the placement deserves, at the length or resolution it needs.

Be clear about what carries over. A different engine makes a different video, not a sharper version of the take you picked, so what you keep from exploring is the direction, not the footage. Judge the new render on its own before it ships.

This is the same logic that makes creative testing work at all, and it fails for the same reason testing usually fails: when each attempt feels expensive, you stop attempting, and you ship your first guess. A cheap engine is not valuable because it saves money on the videos you keep. It is valuable because it changes how many you are willing to throw away.

The reason we are unsentimental about this is that Riffkit runs both engines behind one form. Same source, same creative direction, same product: the engine is a dropdown, and the rate for whatever pair you pick is shown before you render. Moving a direction from one engine to the other is one field in the same form, not a second workflow on a second platform. See what the rates work out to on pricing, or start with a winning video and riff it.

FAQ

Does Seedance 2.5 bill you for the reference video you send in?

Yes. The token count is (input video duration + output video duration) x output width x output height x output frame rate / 1024, so the clip you hand in sits inside the meter alongside the clip you get back. Published guides contradict each other on this point, and our own invoices settle it: a render given a reference video is billed on both. The per-token rate is lower when a video reference is present, which is what makes the mistake easy, but the lower rate only wins if the reference is shorter than about two thirds of the output. Hand a thirty-second reference to a ten-second render and the discount is long gone.

Is MiniMax H3 or Seedance 2.5 cheaper for TikTok ad creative?

H3 is materially cheaper per second of finished video, which is why it is the sensible engine for exploration. It renders up to fifteen seconds per call, which covers most short-form ad creative outright. Seedance 2.5 costs more per second (on Riffkit, 75 credits a second at 480p and 150 at 720p, against 40 for H3 at 768P) and returns up to thirty seconds as one continuous take. For a nine to fifteen second ad variant the premium buys you very little; past fifteen seconds, when the alternative is joining two renders and hiding the seam, it starts to earn its price.

Why does the second segment of a long AI video cost more than the first?

Because handing the previous clip back to the model as a reference raises the token count faster than the cheaper reference-video rate brings it down. Measured on our own production traffic at matched resolution, continuation segments consume about 2.28 times the tokens of a first segment while billing at a per-token rate about 0.6 times as high, which nets out to roughly 1.39 times the cost per second. The intuition that a continuation is cheap because its rate card is cheaper is backwards.

Does changing a finished video's aspect ratio cost the same as making it?

No, and no public comparison covers it. Recutting a finished render into another ratio replays the existing footage through a model, so on an engine that meters the video it is given, an extra ratio bills above that engine's normal per-second rate. On Riffkit an extra ratio is 120 credits a second on Seedance 2.0 at 720p against 100 for the first ratio, 180 on Seedance 2.5 at 720p against 150, and 80 on MiniMax H3 at 768P against 40. If you know you need several ratios, count every one of them in the budget before you render.

Does Seedance 2.5 output 4K?

The official Seedance 2.5 announcement, dated 31 July 2026, does not name an output resolution at all, and the production API we render on offers 480p and 720p only. The 4K figure that circulates in write-ups did not come from that announcement, and we have never seen a 4K tier on the endpoint. The announcement is likewise silent on how many languages the model speaks, so any precise lip-sync language tally you read came from somewhere other than the model's own release post. When a specification decides your budget, check it against the announcement and the API rather than a listing page.

Keep reading

Can Claude Make Videos? What Opus 5.5 Can Do (2026)

Claude does not render video, even Opus 5.5. Three ways people make videos with it, what each is good at, and where realistic AI people come from.

What Counts as "One Video" on an AI Video Pricing Page?

A "video" on an AI video pricing page might be a 5-second clip, a 15-second ad or a talking-actor minute that rounds up. Here is what the main tools count as one, and how to convert any plan into seconds of finished video before you compare.

Write What Changes, Not a Prompt: How to Direct an AI Ad Video Remake

When you remake a winning video with AI, the source already carries the hook, pacing and shots. Write only what should change, placed at the beat where it changes, and fewer renders come back as misses.