How to Keep One Character Consistent Across an AI Video
To hold a character, product or set steady across an AI video, attach several reference images instead of one starting frame — the input mode called reference-to-video. Here are the real per-model limits, why some models need you to name references as @Image1 in the prompt, and why a YouTube link never works as a reference.

To keep the same character, product or set looking consistent across an AI-generated video, you attach reference images rather than a single starting frame. This is a distinct input mode — usually called reference-to-video — and it is not the same thing as image-to-video, which most tools default to. Image-to-video animates one still forward from frame zero. Reference-to-video treats every image you attach as guidance for the whole take, so a face, a garment, a bottle or a colour palette survives across the entire clip instead of drifting after the first second.
The three input modes, and why the difference matters
Almost every complaint about "the AI changed my character halfway through" comes from using the wrong one of these three.
| Mode | What you give it | What it does |
|---|---|---|
| Text-to-video | A prompt only | Invents everything. Nothing is held constant between two runs. |
| Image-to-video | One image | Uses your image as the first frame and animates forward from it. |
| Reference-to-video | Several images, and on some models video and audio too | Uses every attachment as guidance across the whole take. |
The practical test: if you want your product to appear in a scene the model builds around it, you want reference-to-video. If you want your photograph to literally be the opening frame, you want image-to-video.
What you can attach, per model
The caps differ between models, and they differ between versions of the same model — which is the single most common cause of a rejected job. These are the real limits, not rounded ones.
| Model | Reference images | Reference videos | Reference audio |
|---|---|---|---|
| Seedance 2.5 | 30 | 10 (each 1.8 to 30.2 seconds) | 10 |
| Seedance 2.0 and 2.0 Fast | 9 | 3 (each 1.8 to 30.2 seconds) | 3 |
| Gemini Omni Flash 1.1 | 10 | 3, each at most 3 seconds | Not supported |
| Gemini Omni Flash 1.0 | 10 | Not supported | Not supported |
Two of those numbers catch people out. Seedance 2.0 accepts nine images, not ten, so a workflow built around a flat cap of ten will fail on exactly that one model. And Gemini Omni Flash 1.1 caps each reference video at three seconds, for a total of nine seconds across all three — it is designed for short guidance clips, not for feeding in a whole scene.
On Seedance, reference audio is only accepted alongside at least one image or video. Audio on its own is refused.
Name your references in the prompt
This is the step almost nobody documents, and skipping it produces a video that looks like it ignored your uploads entirely.
Seedance addresses references positionally, from the prompt text. Your first image is @Image1, your first reference video is @Video1, your first audio clip is @Audio1. A prompt that never mentions them may legitimately produce a video that appears to have discarded them — that is the model working as documented, not a failure. So write:
"@Image1 walks through the workshop and picks up the tool from @Image2, warm afternoon light, handheld camera."
The Gemini Omni Flash family works differently: reference media is supplied in list order ahead of the prompt, so the order you attach things in is what matters, and you do not name them inline.
Links that look like videos but are not
You can supply a reference by uploading a file or by pasting a link. If you paste a link, it has to be a link that is the file — an MP4 or MOV for video, an MP3 or WAV for audio, a real image file for images.
A YouTube URL does not work, and neither do Vimeo, TikTok, Instagram, Facebook, X or Dailymotion links. Nor do Google Drive and Dropbox share pages. All of those serve an HTML player page rather than the media itself, so what arrives is a web page, not a video. The same applies to any "watch" or "preview" URL. If you can only get a share link, download the file first and upload it.
What this does not do
Three honest limits worth knowing before you plan around this.
- It is guidance, not compositing. Reference images strongly constrain identity, style and palette. They do not guarantee a pixel-exact match, and they are not a replacement for editing a real product shot into a frame.
- On Seedance, your reference video cannot be longer than your output. The combined length of attached reference videos has to stay within the duration of the clip you are generating, because reference length is itself a billable input on that family.
- Some reference images are refused outright. Models run their own content checks on reference media, and images containing identifiable real people are commonly rejected — a rule that applies per model, so the same upload can be accepted by one and refused by another. When that happens you are told which input was refused and no credits are charged for it. Retrying the same image will not change the answer; changing the image will.
What it costs
Reference-to-video is billed the same way as any other generation on that model: per second of output. Seedance 2.5 runs at 13 credits per second of finished video, so a 5-second clip is 65 credits. Gemini Omni Flash 1.1 runs at 5 credits per second, so a 10-second clip is 50 credits. Attaching references does not add a separate fee.
Both of those models sit on the Pro plan. Gemini Omni Flash 1.0, which accepts up to 10 reference images, is available on the free tier — 20 one-time credits, no card — which makes it the cheapest way to find out whether reference images solve your consistency problem at all before committing to a longer render. Current plan details are on the pricing page.
A workflow that holds up
Start with three to five reference images of the same subject from different angles, under similar lighting. Add more only if identity is still drifting; thirty near-identical images is not thirty times better than five varied ones. Name each one in the prompt if you are on Seedance. Generate a short clip first — four or five seconds — confirm the subject holds, then re-run at full length with the same references and the same prompt structure. Changing the reference set and the prompt at the same time makes it impossible to tell which one fixed it.