A prompt guide built from ByteDance’s four official Seedance guides.

Why CineV is writing about Seedance
CineV is built to be the generative AI platform that makes animation best. Read that sentence sideways and a second meaning falls out.
The official Seedance documentation carries this warning: Dreamina Seedance 2.5 and Dreamina Seedance 2.0 series models do not support direct uploads of reference images or videos containing real human faces. Photographs and footage of real human faces cannot be used as reference material. Anyone planning character-driven work hits that wall on day one.
But turn the restriction over and look at the other side. What if the characters are not real human faces at all? What if they are generated animation characters, which is exactly what AI animation work runs on? CineV’s identity as an animation-first AI creative platform starts precisely here, and it is what lets you reach the full range of what Seedance can do.
Add CineV’s 3D-based intuitive conditioning and a second knot comes loose. That maddening experience where the image in your head is perfectly clear and no combination of words will bring it out. Getting past that is a mandatory checkpoint on the way to Seedance’s real capability, and it is where most people give up.
So is this article just a CineV pitch? No. Before any of the platform talk, this article exists to help you understand the Seedance family itself.
What this article is built on
There is no way for anyone outside ByteDance to verify how much the training data and alignment methods changed from version to version. So we started from an assumption. If a single team has been improving one model family toward one goal, then the prompt guides that team published should carry a common pattern that survives across versions.
This article is what happened when we tested that assumption against the official documentation from 1.0 through 2.5. We read all four prompt guides and all three tutorials ByteDance published on BytePlus ModelArk, and every sentence quoted here comes from inside them. On top of that sits what the CineV Creative Team has learned in production.
We hope you reach Seedance 2.5 with the fewest failed generations possible.
The premise: the video has to exist in your head first
Understanding anything after Seedance 2.0 requires one premise up front. To use this model to its limit, the video has to exist in your head before it exists anywhere else.
That will sound like a chicken-and-egg problem. It is one. If the sequence you want has not resolved into an image in your mind, there is no way to give a clear instruction about it. And when the instruction is vague, the model fills the gap on its own terms. What comes back is not what you wanted.
This is not a matter of taste. The official 2.5 guide opens its prompting principles with this line.
Treat Seedance 2.5 as a visual content producer, and write structured prompts with a visual storytelling mindset.
Treat it as a producer, in other words. Producers need a brief. If this is a model that performs at its best when given clear direction and defined roles, then the person who has to supply that direction is you.
If the image is not there yet
CineV’s storyboard feature exists for exactly this gap. Give it a one-line logline and it produces the highlight frames of the entire sequence in order. Think of it as a way to create context when you have none. If you want to unlock Seedance’s full potential but the image in your head is still blurry, start by shaping the scene with CineV Storyboard.
The module that translates the image into the model’s language
Having the image does not immediately produce a prompt. Something has to sit in between and translate.
What you are holding is one of three things: a finished frame, a sequence you can put into sentences, or a pile of scattered material. Each one enters through a different door.

Each stage in turn.
Keyframe reference: the frame that must survive intact
Suppose you already have a finished frame in hand. The character, the style, the acting, the lighting, the camera: all of it works. In that case you pin that frame directly into the output. The instrument for this is Keyframe reference.
Generated visuals align with the keyframes: Input multiple independent images as keyframes, which may include first-frame or last-frame images. The generated video visuals will be relatively strictly aligned with the input images.
To use several frames in sequence, the guide says to write Use Images X to X in order as keyframes. as the first sentence of the prompt. That condition on placement is not decoration. Do not skip past it.
There is a second route to the same destination. Set content.role to first_frame or last_frame. This binds harder, and it charges you for the privilege. The aspect ratio locks along with it. The 2.5 guide states that in this case the ratio parameter must be adaptive and no other value is accepted.
One warning belongs here. A scene image is a unit with the character, style, acting, lighting and camera all fused together. If even one of those elements does not work for you, reconsider forcing that frame. The lighting you dislike gets pinned right alongside the parts you wanted.
Context: the sequence you can put into sentences
Sometimes there is no scene image you like, but the sequence in your head is perfectly clear. The question to ask at that point is single. Can you describe that sequence in the language of the video domain?
If you can, Context is secured. If you cannot, it is not. This distinction matters because starting to load References without Context leaves the model no principle by which to assemble the pieces.
If the translation module is missing
Putting things into domain language is a skill in itself, and not everyone has it. CineV’s Prompt Enhance stands in for that module. It converts the sentences you have into a form the model understands.
Reference: every subject inside the Context
If Context is the sentence, References are its subjects and objects. Seedance 2.5 accepts up to 50 reference assets in a single request: 30 images, 10 videos, 10 audio clips.
The official documentation defines seven kinds.
| Official name | What it references |
|---|---|
Subject reference | Appearance identity and voice of a person, object, scene, or virtual character |
Motion reference | Movement and dynamic information taken from video |
Style reference | The visual style of an image or video |
Audio reference | Music, dialogue, voice, tone, timbre |
Storyboard reference | Subjects, composition, actions, plot, scene progression |
Keyframe reference | One or more images used as keyframes (see the previous section) |
3D clay-model reference/rendering | 3D clay-model video used as a motion reference |
These seven are a functional taxonomy. For planning purposes the grain is slightly wrong. Set them next to the four planning roles from the 2.0 guide, which are far more usable when you are actually assembling assets.
| Planning role | What it does |
|---|---|
Character anchoring | Locks the character’s appearance |
Scene tone-setting | Locks the environment and style |
Camera movement reference | Locks the shot language and rhythm of motion |
Rhythmic atmosphere | Controls emotion and timbre through audio |
Note that backgrounds and props are not broken out separately. That is not an omission. In the official taxonomy both belong to Subject reference. The guide spells it out as “a person, object, scene, or virtual character.”
So how many should you load?
Accepting 50 is not an instruction to supply 50. It is close to the opposite. The 2.0 guide is blunt about it.
It is not recommended to use the full asset limit. Too many assets will make it difficult for the model to judge feature priorities.
The same document gives a recommended configuration as a number. Four to five total. One or two character images, one scene image, one camera-movement video, one audio clip.
The 2.5 recommendations are more granular. Supplying Subject reference as images works well at one to eight subjects; nine to twelve is worth attempting but stability drops and you may need several runs. Supplied as audio or video, one to five works better, with an input duration of five to ten seconds.
You do not need to memorise the numbers. There is one rule. One role per asset. And load only the assets whose role you actually need.
The last check
Once you are here, one thing remains to verify.
Is every subject inside the Context loaded as a Reference, and is each of them called by its own name inside the sentence?
Miss one and the model fills the gap on its own terms. That is what failure actually is.
What actually changed in Seedance 2.5
This is where the real material begins. Before it, one thing needs clearing up. The 2.5 specifications circulating online contradict each other.
Some articles say 2.5 outputs 4K. Others say 4K belongs to 2.0. The companies actually running the API document it as 720p. A reader has no way to know which to believe.
The confusion traces back to a single sentence in the official documentation. Here is how the 2.5 guide states the reference input limits.
Images: Up to 30 images, with resolution up to 4K.
That 4K is the resolution of the input images. Not the output. And the 2.5 tutorial states the opposite about output in plain terms: Seedance 2.5 does not currently support 1080p and 4K resolutions. The platform-wide tutorial confirms that the only model capable of 4K output is Seedance 2.0.
Here are the boundaries the official documentation actually sets for 2.5.
| Item | 2.5 | Source |
|---|---|---|
| Maximum length, single generation | 30 seconds | 2.5 guide, overview |
| Output resolution | 720p (no 1080p, no 4K) | 2.5 tutorial |
| Frame rate | 24 fps | Platform tutorial |
| Reference assets | 50 total (30 image · 10 video · 10 audio) | 2.5 guide |
| Aspect ratio | Free within the range [0.4, 2.5] | 2.5 guide |
| Native languages | More than 10 | 2.5 guide |
So what improved in real terms over 2.0? This does not require guesswork either. The 2.5 guide contains a section called Differences from Seedance 2.0, and it lists exactly four items.
Seedance 2.0 does not respond to timestamps and only responds to shot numbers, while Seedance 2.5 supports integer-second timestamps.Seedance 2.0 does not recommend using multi-view images as subject references, while Seedance 2.5 supports them.Seedance 2.0 only supports six fixed output aspect ratios, while Seedance 2.5 can support any output aspect ratio between [0.4, 2.5] by controlling the input assets.Seedance 2.5 supports MOV output, which better preserves color consistency, brightness consistency, and audio-visual consistency in extension and editing tasks.Seedance 2.5 divides tasks into two categories based on whether the input reference assets lock the properties of the output video. Seedance 2.0 does not make this distinction.
The fifth item appears in a different section of the document, but it belongs with the others. Tasks are now split in two according to whether the input assets lock the properties of the output, a distinction that did not exist in 2.0.
Items one and five take up the rest of this article. Both of them manufacture failures quietly.
The single formula shared by all four official guides
Time to test the assumption we started with. From each official prompt guide, 1.0 through 2.5, pull out only the sentence that states the prompt formula, and set them side by side.
| Version | The prompt formula as the official documentation states it |
|---|---|
| 1.0 pro | Beginner: subject+action |
| 1.5 pro | Subject + Movement + Environment (optional) + Camera movement (optional) + Aesthetic description (optional) + Sound (optional) |
| 2.0 series | precise subject + action details + scene/environment + lighting & color tone + camera movement + visual style + image quality + constraints |
| 2.5 | Subject + Location + Event + Genre/Style + Camera movement... |
The slot count goes from two to six, up to eight, then back down to five. The names shift too. Movement becomes action details, then becomes Event.
But the order is identical in all four.
The subject comes first. What the subject is doing comes second. Where it happens comes third. Camera comes after that. Across two years and four documents, that arrangement never moved once. Whether ByteDance held it deliberately, or whether it emerged because the models kept learning from the same pool of data, is not something we can determine from outside. What we can say is that four repetitions is difficult to read as coincidence.
The practical conclusion follows. Whatever version you are on, when a prompt stops working, return to this order.
Who → doing what → where → and how the camera behaves
Version 2.5 wraps one more layer around it. The formula above is only the first block in 2.5, where it is called One-Sentence Summary. The full structure has four blocks.
| Block | Role |
|---|---|
Asset Referencing for R2V | Declare what each uploaded asset is responsible for |
One-Sentence Summary | The formula above. The whole piece in one sentence |
Detailed Plot Description | Segment-by-segment description, divided by Shot or by timestamp |
Additional Notes | What should hold constant across every segment |
The guide demonstrates these four blocks with an example of its own. This is the original, verbatim.
Realistic nature documentary style, natural lighting and shadows. On a warm
afternoon, on a grassy slope in the forest, a chubby panda cub rolls down the hill.
The panda has fluffy, realistic black-and-white fur, a small round body, and clumsy,
adorable movements. The scene is a green forest slope. The ground is covered with
grass, moss, clover, soil, small stones, dry branches, and a few small yellow flowers.
Tall tree trunks and dense woods are softly blurred in the background. The camera is a
low-angle medium-wide shot with a slight handheld feel. The framing remains mostly
stable, keeping the panda in frame at all times.
0s-3s: A panda cub lies on a green grassy slope, its body round and chubby. It begins
to slowly roll sideways down the slope with clumsy movements, gently bending the grass
beneath its body. A light breeze passes through, and sunlight filters through the trees
from the upper left, creating dappled light and shadow.
3s-8s: The panda rolls toward the lower right of the frame and gradually comes to a
stop, shifting from lying on its side to lying on its belly. Its round face turns toward
the camera, and its front paws press into the grass. The panda lies in the foreground
grass, adjusts into a comfortable position, slightly raises and lowers its head, and
makes a soft little humming sound.
Low camera position, slight handheld feel, subtly following the panda as it moves toward
the lower right. Natural depth of field: the foreground grass is slightly blurred, the
panda remains clear, and the background forest is softly out of focus. Natural
environmental audio only, including wind, rustling grass, and the soft plop of the panda
rolling. The overall mood is warm, realistic, and natural.
Four paragraphs mapping onto four blocks. The first is the one-sentence summary, the second holds the fixed information, the third and fourth are the segment descriptions, and the last collects what stays constant. And note that this example divides its time into second-based intervals. That is the subject of the next section.
Here is how the guide summarizes the whole structure in one line.
Treat Seedance 2.5 as a visual content producer, and write structured prompts with a visual storytelling mindset.
The same sentence quoted earlier. What it is asking for should be concrete now.
Integer-second timestamps: why “use shot numbers, not timecodes” is now wrong
Anyone who has searched for Seedance prompting advice has run into this rule. Do not pin exact seconds. Label your shots, Shot 1 and Shot 2, and let the model find the pacing.
That advice was never wrong. It was correct through 2.0.
The 2.0 guide states it directly.
Do not impose strict limits on the duration of each segment; prioritize allowing the model to naturally generate the pacing based on the plot.
The same document gives its reasoning: The model's support for precise timing (such as 0–3 seconds) is unstable, and forcibly limiting duration may lead to abnormal generation results. Precise timing was unstable, so do not use it. That was a sensible instruction.
The 2.5 guide says the reverse.
Seedance 2.0 does not respond to timestamps and only responds to shot numbers, while Seedance 2.5 supports integer-second timestamps.
Version 2.0 did not respond to timestamps. Version 2.5 supports them at integer-second resolution. The vendor reversed its own guidance in its own documentation.
The problem is that the market has not absorbed the change. A large share of the Seedance prompting guides currently ranking were written in the 2.0 era, and they still teach you not to use seconds. In some cases a single publisher runs a 2.0 article and a 2.5 article side by side with contradictory instructions and never retracts either. Their 2.5 article demotes timestamps to an occasional tool, on the very version where timestamp control is a headline improvement.
So how should seconds be used in 2.5? The guide names three methods.
Interval control. The default approach, with one condition attached.
Clear time intervals. Pay attention to timeline continuity and avoid gaps such as “0-3s… 5-6s…”.
Do not leave holes between intervals. Chain them: 0-3 seconds … 3-7 seconds … 7-15 seconds.
Point control. Pin an event to a specific moment. “Quick left sideways transition at the 5-second mark.”
Relative control. Anchor to an event rather than to absolute time. “John stands there blankly. After 3 seconds, everyone around him shakes their head.”
The base unit is one second. The guide states Use 1-second intervals as the basic unit. It records the limits as well. Do not try to control high-frequency motion through timestamps, such as shaking a head three times per second. Put too little into an interval and the model improvises freely; pack too much in and the result either fragments into excessive cuts or drops part of the plot.
The summary is this. Use seconds, but do not mistake seconds for frame-accurate edit points. A timestamp allocates time to events. It is not a cut list.
In one line
Shot numbers still work. Version 2.5 accepts both. What no longer applies to 2.5 is the 2.0-era warning that seconds will break your generation.
Declaring References and assigning their roles
Uploading a Reference and having a Reference take effect are two different things. Upload it without naming it in the sentence, and the model has no way to determine what to do with it.
The 2.0 guide supplies a sentence pattern for declaring a subject.
Define [Core_Subject_Features] in <Image/Video_N> as <Subject_N>
A condition comes attached: Use 2–3 clear and stable static features (such as clothing, hairstyle, appearance, or category) to describe the subject and ensure it can be uniquely identified. Two or three stable, static features. Anything that changes, an expression or a gesture, cannot serve as an identifier.
Version 2.5 asks you to spell out the correspondences as a list when several assets are in play. The examples it gives look like this.
Images 1-2 are Character 1 and correspond to Audio 1; Images 3-4 are Character 2 and correspond to Audio 2.
Image 1 depicts the protagonist John and uses the voice timbre from Audio 1.
Refer to Image 1 for lighting and filters.
The third one matters most. It does not reference the image wholesale. It narrows the scope to the lighting and filters only.
When there are multiple characters, label them and use nothing but those labels from then on. The guide’s example is clean.
Define the tall man in Video 1 as police officer, and define the other short man as thief.
For the rest of the prompt that person is always police officer, always thief. Switch to “the man” or “the taller one” and the instruction breaks. The same document carries a worked example applying the principle.
Define the tall man in Video 1 as police officer, and define the other short man as
thief. The scene is set in a crowded daytime market, with bright sunlight, many fruit
stalls, and dense pedestrian traffic, creating a lively street-market atmosphere.
Thief runs forward in panic through the crowded market, while police officer follows
closely behind at full speed. The two quickly weave through the stalls. A handheld
camera rapidly tracks forward, with slight realistic camera shake, creating a tense
chase atmosphere.
Notice that the sentence defining the labels and the sentences using them sit inside one continuous prompt. There is no separate setup block.
Order carries meaning. The 2.0 guide states Place important assets first: The more an asset requires precise reference, the earlier it should be placed in the prompt. Numbering follows upload order, which makes the upload order itself a declaration of priority.
Version 2.5 adds two more requirements. First, state what you are referencing it for. “Refer to the action of casting the spell in Video 1 and the wrap-around camera movement in Video 2.” Second, when the asset is already accurate enough, stop describing the scene again. Adding “he raises a hand, he turns around” on top becomes interference rather than instruction.
Here is the official example for pulling motion out of a video.
Refer to the character movements and shot language in Video 1 to create a fight scene
with the character from Image 2 on the left and the character from Image 1 on the right.
Include intense background music.
What to take from Video 1 is narrowed to character movements and shot language, and Image 1 and Image 2 are assigned left and right positions. Reference scope and blocking, in one sentence.
The documentation is equally explicit about what not to do.
It is not recommended to provide mapping information only inside the image itself.
Writing a name onto a character image and then referring only to that name in the prompt. It causes character confusion and duplicate generation.
The official documentation could not standardize its own notation
Something needs stating plainly here. The notation the 2.5 guide prescribes is Image 1 / Video 1 / Audio 1. And yet four different notations appear across the official examples in that same document.
| Notation | Where it appears |
|---|---|
Image 1 | The prescription itself, and most examples |
@Image 1 | Editing and extension examples |
@image1 | Locked-task trigger examples |
[Image 1] | 3D clay-model examples, asset binding lines |
<1pic> … <10pic> | Detailed 3D clay-model example |
There is no basis for declaring any one of them correct. A working principle does follow, though. Pick one and hold it for the length of a prompt. The moment you mix notations, you give the model room to treat one asset as two.
Bracket syntax for audio, dialogue and subtitles
The 2.0 guide provides a convention for separating information types by punctuation, presented as a table.
| Information type | Symbol | Example |
|---|---|---|
| Music | () | (fast-paced rock music is playing in the background) |
| Sound effect | <> | < dog barking can be heard in the distance > |
| Dialogue | {} | {Hello, world} |
| Subtitles | 【】 | 【Chapter One: Departure】 |
Dialogue carries one condition. For any language other than Chinese or English, the language has to be named: says in Japanese {こんにちは}.
There is something to watch here. The 2.5 guide does not contain this table. The curly-brace notation does not appear anywhere in that document. What 2.5 provides instead is a different instrument.
Supports negative control for subtitles. … Supports negative audio control for finer dimensions, including sound effects, background music (BGM), and dialogue.
Control by negation: Do not add subtitles. or No BGM; generate only environmental sounds and action sounds. The 2.5 guide sets its general rule as Use positive descriptions whenever possible, and permits negative instructions only for subtitles and audio.
So has the bracket syntax been retired in 2.5? There is no basis for saying so. The 2.5 guide notes that subject, motion, audio and style reference usage is unchanged from 2.0 and points to the 2.0 document for detailed examples. What is worth knowing is that the mechanism 2.5 itself vouches for is negative control.
Designing a 30-second single take
The 30 seconds in 2.5 is not 15 seconds doubled. In the guide’s own words the goal is presenting a complete story without multi-segment stitching. That raises the design burden accordingly.
The skeleton is the four blocks already covered. In practice it lays out like this.
[1] One-sentence summary
Subject + Location + Event + Genre/Style + Camera movement
[2] Fixed information
Character appearance, space, base camera setup
[3] Segment description
0-3s: …
3-8s: …
8-15s: …
[4] Hold across all segments
Camera behaviour, lighting, sound, overall mood
Within a segment, the ordering rule from 2.0 still applies. Write camera movement → the subject’s action and expression → change in position → sound, in that order.
Version 2.5 sets separate principles for describing action.
Actions: Give priority to general descriptions … Only write specific details for a few memorable actions, and avoid repeating the same actions.
Expressions: Use descriptive sentences and reduce the use of idioms.
Write action in large strokes and detail only the few moments that matter. For expressions, cut the idioms and write descriptively. The 2.0 guide supplies this principle in a far more concrete table, the point of which is to replace an emotion label with observable physical behaviour. Instead of “sadness”, write lowering the head, shoulders trembling slightly, eyes reddening, fingers unconsciously clutching the corner of clothing.
Camera terminology comes as a list in 2.5. Shot size is extreme wide shot/wide shot/medium shot/medium close-up/close-up, movement is push in/pull out/pan/track/follow/orbit/dive/pull back/tilt up/handheld shake, angle is low angle/overhead shot/first-person perspective. There is a rule for anything outside the list.
For overly niche or technical terms, convert them into [term + descriptive explanation].
Do not throw the term in bare. If you want Rack focus, write where the focus travels from and to alongside it.
Locked and Unlocked: when the request is silently ignored
Now for the distinction 2.5 introduced. This is the one that produces the most frequent and least detectable failures in production.
The symptoms usually present like this. You specified ratio for a 16:9 output and got back a vertical video. You asked for ten seconds and got seven. You set the parameter, no error came back. And nothing you specified took effect.
The cause is the task type. Here is how the 2.5 guide explains it.
Seedance 2.5 divides tasks into two categories based on whether the input reference assets lock the properties of the output video. Seedance 2.0 does not make this distinction.
Locked means the input asset is laid onto the output timeline directly. The model conforms to the input, so the aspect ratio locks, and in some cases the duration with it. Unlocked means the input serves only as a semantic reference, leaving you free to set aspect ratio and duration.
Three tasks fall under Locked.
| Task | What locks | Required value |
|---|---|---|
| Video editing | Aspect ratio and duration | ratio=adaptive, duration=-1 |
| First / last frame | Aspect ratio | ratio=adaptive (duration is yours) |
| Video extension | Aspect ratio | ratio=adaptive (duration is yours) |
For editing, the duration locks too, and it does not come back identical. The guide notes that the output duration may differ slightly from the input, by up to about 0.3 seconds, explaining that some transition frames get compressed by the frame-handling mechanism. Feed a video that 2.5 generated back in as editing input and this discrepancy disappears.
There is a harder part. The task type is determined by the words in your prompt. To register as editing, the prompt has to contain a trigger: edit video, add, insert, remove, delete, modify, replace, change to. For extension the triggers are extend forward, extend backward, continue, continue from, extend the story.
Use the wrong words and the model misclassifies the task, which means a different set of parameter rules gets applied than the one you expected. When several videos are uploaded, the model determines which video to edit based on the prompt. Even the choice of which video is decided by your wording.
The guide pairs its principle for writing edit instructions with examples. The principle is state what changes from what to what.
Change the man's action from drinking coffee to mopping the floor from 4-6 seconds in Video 1, and leave the rest of the content unchanged.
Editing task: Replace the Asian woman on the right in Video 1 with the Latina woman from Image 1.
The first specifies an interval and adds a sentence telling the model to leave everything else alone. The second opens with Editing task: and pins the task type at the front, a safety measure to guarantee the trigger word lands.
Editing and extension carry a format recommendation as well: It is recommended to set output_format to mov. The reason given is better preservation of colour, brightness and audio-visual continuity.
Remember it by the symptom
Aspect ratio changed on its own: you were in a Locked task. Duration off by 0.3 seconds: the editing task worked correctly. Your edit instruction ignored entirely: no trigger word, so it never registered as editing.
Prompts for animation work
Everything above comes out of the official documentation. Now to translate it into the form the CineV Creative Team actually uses for animation. The examples below follow the official principles, but the sentences are ours.
Holding one character across 30 seconds
Use Images 1 to 3 in order as keyframes.
A hooded courier in a rain-soaked neon alley delivers a sealed package,
2D cel-shaded anime, slow push-in.
Subject 1: the courier in a mustard-yellow hooded jacket with a frayed
left sleeve and a cracked visor. Refer to Image 4 for the character design.
Environment refers to Image 5 for the alley's color and signage.
0-6s: Subject 1 walks toward camera through standing water. Camera pushes
in slowly from a wide shot. Reflected signage ripples underfoot.
6-14s: Subject 1 stops, looks down at the package, and adjusts the strap.
Medium shot, camera holds.
14-24s: Subject 1 raises the visor. Close-up. The face stays consistent
with Image 4.
24-30s: Subject 1 turns away and walks into the alley depth. Camera pulls
back to a wide shot.
Throughout: 2D cel-shaded anime with hand-painted backgrounds, cool cyan
and magenta palette, rain ambience and distant traffic only. No BGM.
No subtitles. Keep the character design consistent with Image 4 in every
shot.
Four things to check in that prompt. The keyframe declaration is the first sentence. The subject is identified by two or three static features. The intervals chain without gaps. And the closing block collects everything that holds constant. All four come straight from the official principles quoted above.
Taking motion from another video
Refer to the fight choreography and camera work in Video 1 to generate a
rooftop duel between Subject 1 and Subject 2.
Subject 1: the silver-haired swordsman in a black high-collar coat,
refer to Image 1.
Subject 2: the masked opponent in red armor plating, refer to Image 2.
Refer to Image 3 for the rooftop environment and its lighting.
Keep the motion timing consistent with Video 1. 2D anime with strong
key-light rim separation.
The Refer to … in Video 1 construction is the Motion reference template from the 2.0 guide. Narrowing what to reference down to fight choreography and camera work follows the 2.5 requirement.
For animation work the restriction discussed at the top turns into an advantage. Real human faces being blocked from upload means that using generated character designs as References is the intended path, not a workaround. It is the first of the three routes the official documentation itself proposes.
What to check when the result misses
Failures tend to have specific causes. The symptoms covered in the 2.0 guide’s FAQ are the ones encountered most often in production, and the following rearranges its prescriptions into a checking order.
The character’s face changes partway through. The 2.0 guide calls this ID drift. The prescription is to prepare a separate close-up of the face: front-facing, no expression, with interference from shoulders, neck and background kept to a minimum. A face close-up and a full-body shot are sufficient for a character Reference; three-view and multi-view images are not recommended.
The same character appears twice in frame. The guide names this the twin problem and supplies a global constraint to append at the end of the prompt. Throughout the video, characters with completely identical appearance, clothing, and accessories are prohibited. Do not generate duplicate avatars or a twin effect. It also gives a number: stability degrades past four characters.
The style drifts partway. Pin the style constraint explicitly. Write out whether it is 2D Japanese animation style or 3D. For tighter control, convert the Reference image itself into the target style before feeding it in.
Unwanted subtitles appear. Add Keep it subtitle-free or Avoid generating any text or subtitles. For 2.5, Do not add subtitles. is the official form. Worth noting: the 2.0 guide records that vertical framing produces subtitles at a significantly higher rate than horizontal.
Quality degrades with each extension. The guide’s advice is not to stack extensions repeatedly. Use a high-resolution original as the Reference and keep the extension count under control.
Before any of this, check one thing. Did you load too many assets? The 2.0 guide recommends four to five in total. Accepting 50 is a ceiling, not a target.
Using Seedance 2.5 on CineV
That covers what reading the Seedance family can give you. What remains is carrying these principles into actual work, and there are two places where that process reliably stalls.
The first is having no image in your head, and therefore no Context to build. The second is having the image but being unable to put it into domain language. Storyboard and Prompt Enhance were mentioned earlier for each. They handle the front and back ends of the translation module.
On top of that sits the identity of an animation-first platform. Working along a path where the real-human-face restriction never bites, and using 3D-based intuitive conditioning to handle the image that will not resolve into words. Which is where this article started.
We hope you reach Seedance 2.5 with as few failed generations as possible. CineV is currently offering 50% off.
Frequently asked questions
Can I use second-based timestamps in a Seedance 2.5 prompt? Yes. The 2.5 guide states support for integer-second timestamps. The base unit is one second, and the condition is that intervals must not leave gaps between them.
Do “Shot 1, Shot 2” labels still work in 2.5? They do. Version 2.5 accepts both shot numbers and timestamps. What does not carry over is the 2.0 warning that using seconds destabilises the result.
Does Seedance 2.5 output 4K? No. The 2.5 tutorial states that 1080p and 4K are not supported. The 4K figure refers to the maximum resolution of input images; the model capable of 4K output is Seedance 2.0.
How long can a single generation be? Up to 30 seconds.
Can I change the aspect ratio and duration when editing a video? No. Editing is a Locked task, so ratio must be adaptive and duration must be -1. The output length can differ from the input by up to about 0.3 seconds.
How many References can I use? Fifty in total for 2.5: 30 images, 10 videos, 10 audio clips. The 2.0 guide recommends four to five and warns that filling the ceiling makes it difficult for the model to judge priorities.
Can I use a photograph of a real person as a Reference? No. The official documentation states that direct upload of images and videos containing real human faces is not supported. It offers three alternatives: model-generated output, pre-made digital characters, and material that has been rights-cleared.
Where can I read the official Seedance 2.5 prompt guide? It is published in the BytePlus ModelArk documentation. Every source is linked below.
Official sources
Every quotation in this article comes from the documents below.
- Seedance-1.0-pro&pro-fast prompt guide (BytePlus ModelArk)
- Seedance-1.5-pro prompt guide (BytePlus ModelArk)
- Dreamina Seedance 2.0 series prompt guide (BytePlus ModelArk)
- Dreamina Seedance 2.5 prompt guide (BytePlus ModelArk)
- Dreamina Seedance 2.0 series tutorial (BytePlus ModelArk)
- Dreamina Seedance 2.5 tutorial (BytePlus ModelArk)
- Video generation tutorial (BytePlus ModelArk)
The documentation is updated continuously. This article reflects the state of it as of August 2026. Rights to the quoted passages rest with the publisher of each document.

