← Blog

Write What Changes, Not a Prompt: How to Direct an AI Ad Video Remake

Most advice on AI video prompts assumes you start from a blank page: describe the creator, the setting, the hook, the tone, the call to action, and hope the model assembles something that works. When you remake a video that already won, most of that work is done before you type anything. The source already carries the hook, the pacing, the shot order and what lands on the beat. Your job is not to describe a video. It is to describe the difference.

That matters for a plain reason: every successful render is billed by the second. A vague direction does not fail loudly. It comes back as a video that looks fine and does not land, and the fix is another render. This guide is about writing the direction so fewer takes miss.

The engine sets the ceiling, your writing sets the take

The same video model has the same ceiling in any app that calls it: how sharp and lifelike the frames can get is set by the engine, not by the tool around it. So when a take comes back unusable, the cause is usually in the request, not the engine. Within that ceiling, how good a take gets depends mostly on what you wrote. Takes that miss are usually vague, stuffed with several ideas, or asking the model for something it renders badly.

All three are fixable before you spend anything.

Start from the source, then write only what changes

When you riff a winning video, everything you leave out is inherited: shot sizes, beat order, pacing, where the music hits. That is the whole point of starting from something proven, so treat the source as the draft and your direction (the Creative direction box when you Adapt) as the edit.

A useful test: if the only thing changing is who is on camera, leave the direction empty. Picking a character already does that. Writing more pulls the render away from a structure that already works.

Compare two directions for the same source, a "restock my fridge" routine remade for an insulated bottle:

  • Weak: Make an engaging UGC ad for our insulated bottle that shows how great it is and appeals to gym-goers.
  • Strong: Same restock routine, but it's a gym bag, not a fridge. As the last item goes in, she drops the bottle in and zips the bag. Caption at that beat reads "ice at hour nine".

The weak one restates what every ad wants and gives the model nothing it can place. The strong one names one change, says where it happens, and fixes the one line that must read exactly.

Put every change at a beat

A change stated as a concept loses to the source. "Make it about hydration" gets diluted, because the source's own content is stronger than an abstract request. The same change stated with a place (which beat, what happens right before and right after) is the one that shows up on screen.

A beat is one moment of the video: the hook, the reveal, the payoff. To see them, open the template in Templates: the Video analysis section on its Detail tab lists each beat with its time. Then write your changes against them: "at the reveal", "on the second cut", "in the closing line".

One video, one point

The fastest way to ruin a take is to list every selling point. Six features in fifteen seconds read as a feature tour, and a viewer who scrolls past keeps none of them. Pick the one thing this video is for. The other five are the next five videos, each built on the same winning structure with a different angle.

Length follows how far you depart, not how much you care

A small change needs a line. A big departure, like moving the whole routine into a different room or turning a talking-head into a demo, needs detail, because the model has to invent more. Writing three paragraphs because this video matters to you does not help; it adds instructions the model has to weigh against a source that was already working.

Keep "how it's shot" apart from "what's in it"

How it's shot (light, grain, camera feel) and what's in it (wardrobe, props, setting) are two different things. Put them in one sentence and one drags the other: asking for an unpolished look often flattens the subject too. Give the look its own sentence.

One more thing about the look: if you describe it, yours replaces the source's. If you say nothing, the remake follows the source's look, and every source looks different. If you post several videos to one account, decide the look once ("handheld phone, uneven daylight, no beauty filter, visible skin texture") and reuse that sentence every time, so the account's look doesn't shift with every source.

Quote the words that must not change

Put any line or caption that must read exactly in double quotes. Quoted text is kept word for word and never translated, even when the video's language is different: she says "Don't overthink it." keeps that English line inside a Spanish video. Everything unquoted is direction, and gets written naturally in the video's language.

Don't ask the model for what it renders badly

Some requests fail no matter how well you write them, because they are structural weaknesses of video models:

  • Screens inside the video. Phones, app interfaces, a video playing inside the video. They come out unconvincing. Show the result as a real scene instead: the person in the kitchen, not a phone showing the kitchen.
  • Text in the scene. Signs, shirt prints, a printed end card. The model invents letterforms and garbles them. Captions are a separate layer, burned in with a real text renderer, so put your claim in the caption. We explain why in why AI video garbles text.
  • Your product's label. This one works far more often, because the product is supplied as an image, though small print can still break. Upload a large, sharp photo, give it a name on the product page (say "front label"), set placement to On-camera, and write that name in your direction at the beat where it appears (at the reveal she turns the front label to camera). Typing @ in the box picks it for you. Once you name an image, only the images you name are used in that video. Let the label be the subject of the shot, not the background.

Bringing a new video? Say what made it pop

When you riff a new video (a link or an upload), it is analyzed first, and that read is accurate on what is visible and weaker on what is implied. We looked at 621 human corrections to model readings of viral videos: almost none were about objects. They were about relations. Who is who, what is locked to the music, what a cut caused, what is wrong on purpose.

So if the source's secret is the twist at 0:03 or the opening line, say so in one sentence in the What made it pop? box under Analyze a new video: "the same person plays both roles", "every outfit change lands on the drum hit". The analysis reads the source with that in mind. For a template you added yourself, open it in Templates, choose Re-analyze and put the same sentence under Viral insight.

Test the angle on the cheaper engine

Engines cost different amounts per second. MiniMax H3 at 768P is 40 credits a second, Seedance 2.0 at 720p is 100. A 15-second test uses 600 credits on H3 and 1,500 on Seedance 2.0.

Use that gap to test the idea, not the polish. Run an angle and its lines on H3; if it works, render the version you post on Seedance 2.0 from the same direction (Seedance needs a paid plan). Keep in mind the Seedance render is a new take, not the H3 video upscaled, so check it too. The real costs of each engine are broken down separately.

A run that produces no video costs nothing; if a longer video fails partway, only the seconds already rendered are billed. Everything that renders is billed by the second, which is why everything above happens before you submit.

Keep a guard list

When a take comes back with something you never asked for (a studio key light, a smiling reaction where you wanted deadpan, an extra person in frame), that is the model's default showing. Add an explicit "not X" to your next direction. People who riff often keep a short list of these and paste it into every direction. It is the cheapest improvement there is.

Before you submit

Read your direction once against this list:

  • Does it say only what differs from the source?
  • Is every change placed at a beat?
  • Is it one point, not a tour?
  • Is the look in its own sentence, and the same one you used last time?
  • Are the exact lines in double quotes?
  • Did you avoid screens and in-scene text?

If you want the character side of this, keeping the same AI character across videos covers what to lock. Then open a template and write your first direction, or let an agent draft it with you through the Riffkit skill.

FAQ

How do I write a prompt for an AI UGC ad video?

If you start from a video that already performs, don't write a full prompt. The source already carries the hook, pacing, shot order and beat timing, so write only what should differ from it: one change per beat, placed where it happens ("at the reveal she holds the box to camera"). Put any line or caption that must read exactly in double quotes, keep the visual look in its own sentence, and make the video about one point instead of a list of features.

How long should an AI video prompt be?

As long as the change you are asking for, not as long as the video is important. When you remake a winning video, a small change such as a different product or a new closing line needs one or two sentences, and if only the person on camera changes you can leave the direction empty. A big departure, like moving the routine into a new setting or turning a talking-head into a demo, needs more detail because the model has more to invent.

Why did my AI video come out nothing like my prompt?

Usually one of three things: the request was abstract, so the source's own content won out; the prompt stacked several ideas, so none of them got enough screen time; or it asked for something video models render badly, such as screens, app interfaces or text on signs and clothing. Placing each change at a specific beat, keeping one point per video and moving words into captions avoids most of these before the next render.

How do I keep an exact line or caption in an AI video?

Put it in double quotes. On Riffkit, quoted text in the creative direction is kept word for word and is not translated, even when the video is generated in another language, so a slogan in English stays in English inside a Spanish video. Unquoted text is treated as direction and is written naturally in the video's language.

Can AI video models render text, phone screens or product labels?

Text inside the scene is their weakest point: signs, shirt prints and printed end cards usually come out garbled, and phones or app screens inside the video look unconvincing. Captions avoid the problem when they are burned in as a separate layer after generation, which is how Riffkit adds them. Product labels work far more often because the product is supplied as a reference image, though small print can still break: upload a large, sharp photo, make the label the subject of the shot, and check the finished video before you spend on it.

Should I test AI video ideas on a cheaper engine first?

It is a sensible way to test an angle and its lines before spending on the final version. On Riffkit, a new riff on MiniMax H3 at 768P costs 40 credits a second and Seedance 2.0 at 720p costs 100, so a 15-second test uses 600 credits instead of 1,500. When the idea works, render it on Seedance 2.0 (paid plans) from the same direction. That is a new take rather than an upscale of the test, so review it before you post.

Do I pay for AI video renders I don't end up using?

On Riffkit, every successful render is billed by the second at the engine's rate, whether or not you post it. A run that produces no video costs nothing, and if a longer video fails partway, only the seconds already rendered are billed. That is why the direction is worth getting right before you submit: a vague one rarely fails outright, it comes back as a usable-looking video that doesn't land, and the fix is another render.

Keep reading

Can Claude Make Videos? What Opus 5.5 Can Do (2026)

Claude does not render video, even Opus 5.5. Three ways people make videos with it, what each is good at, and where realistic AI people come from.

Higgsfield Genjutsu Alternative: Other Ways to Reshoot a Video (2026)

Looking for a Higgsfield Genjutsu alternative? What Genjutsu does well, how face swap and character swap and full reshoots actually differ, and when remaking the shot in your own pipeline fits better.

How to Use Higgsfield Genjutsu: A Real Guide to Reshooting a Video (2026)

How to use Higgsfield Genjutsu: Motion Transfer vs Object Swap, the 4 to 30 second limit, reference images, how credits work, and where reshooting a video gets frustrating.