Blog
Guide de prompts Seedance 2.5 : réaliser des vidéos de 30 secondes comme un vrai réalisateurGUIDE
11 août 202613 min de lecture

Guide de prompts Seedance 2.5 : réaliser des vidéos de 30 secondes comme un vrai réalisateur

Seedance 2.5 changed what a single generation can be. One clip can now run a full 30 seconds in one continuous piece. You can hand it up to 50 reference materials at once — images, video, audio — and it will keep track of who is who. It speaks more than ten languages. And it can edit or extend a video you already have instead of forcing you to start over.

All of that power only shows up if you write for it. A one-line prompt gets you a one-line result. The creators getting the best output from Seedance 2.5 are not writing better sentences — they are writing like directors: assigning roles, laying out a timeline, calling the shots.

This guide walks through exactly how to do that.

What You'll Learn

  • The four-part prompt structure that works for almost every Seedance 2.5 job
  • How to map your reference materials so the model never confuses two characters
  • How to use timestamps to control story beats second by second
  • Camera language, action, and expression that the model actually understands
  • Task-specific playbooks: keyframes, storyboards, blockouts, editing, extending, montage
  • The mistakes that quietly ruin otherwise good prompts

Think Like a Director, Not a Search Engine

The single biggest mindset shift: stop describing a picture, start directing a shoot.

A search-engine prompt sounds like this:

a panda rolling down a hill

A director's prompt sounds like this:

Realistic nature-documentary style, cinematic natural lighting, a warm afternoon on a forest slope. A round panda cub tumbles down the grass.

Same idea. Completely different result. The second version already tells the model the genre, the light, the time of day, the location, and the subject's physical character — before a single second of motion is described.

Seedance 2.5 rewards that specificity everywhere, and the structure below is how you deliver it consistently.

The Four-Part Prompt Structure

Almost every strong Seedance 2.5 prompt has the same four layers, in this order:

1. Material mapping   → which reference is what
2. One-line summary   → subject + place + event + genre/style
3. Timeline detail    → beat by beat: action, camera, dialogue, sound
4. Global closing     → what stays true across the whole clip

Here's a complete example with no references at all — pure text to video:

Realistic nature-documentary style, cinematic natural lighting, a warm afternoon
on a forest slope, where a round panda cub tumbles down the hill.

The panda has fluffy, realistic black-and-white fur, small and chubby, with
clumsy, endearing movement. The setting is a green forest slope covered in grass,
moss, clover, soil, small stones, fallen twigs, and a few small yellow flowers,
with tall trunks and woodland blurred in the background. Low camera position,
medium-wide framing, a slight handheld feel, mostly static, keeping the panda in
frame throughout.

0s-3s: The panda cub lies on the green slope, body round, and begins to roll
slowly sideways down the incline. Its movement is clumsy; blades of grass bend
softly under its weight. A light breeze passes through, and sunlight filters from
the upper left through the trees, casting dappled light.

3s-8s: The cub rolls toward the lower right of the frame and gradually comes to a
stop, shifting from lying on its side to lying on its belly. Its round face turns
toward the camera, front paws pressed into the grass. It settles into the
foreground undergrowth, adjusts to a comfortable position, lifts its head slightly
and lowers it again, letting out a soft little grunt.

Low camera position, slight handheld feel, drifting gently with the panda toward
the lower right. Natural depth of field: foreground grass slightly soft, the panda
sharp, the woodland behind softly blurred. Natural ambient sound — wind, and the
soft thump of the panda rolling. Overall warm, real, and natural.

Notice what the closing paragraph does. It doesn't add new events — it locks in the things that must hold across all 30 seconds: camera behavior, depth of field, ambience, mood. That final layer is what keeps a long clip from drifting.

Mapping Your References: The Rule That Prevents Chaos

With up to 50 reference materials in play, the mapping is more important than the description. If you upload six images and simply write "make them fight," you're gambling.

References are numbered in the order you upload them — image 1, image 2, video 1, audio 1, and so on. Bind each one to a role in the text.

Good mapping looks like this:

Images 1-2 are Character A, who uses audio 1 for their voice. Images 3-4 are Character B, who uses audio 2.

Image 1 is the lead, Zhang San. Image 1 uses the voice from audio 1.

Do not put the mapping inside the images. Writing a character's name on the reference picture and then referring to that name in your prompt is one of the most common ways to end up with duplicated or merged characters. Say it in the text.

Say what to take from each reference

A reference is rarely something you want copied wholesale. Be specific about which part matters:

Reference the spellcasting motion from video 1, and the orbiting camera move from video 2.

Reference the lighting and color grade from image 1.

When a reference is precise, stop describing

This one is counterintuitive. If your reference video already contains exactly the movement you want, don't narrate it again — restating it in words gives the model a second, competing instruction.

Strictly follow the motion and camera work in video 1, keeping the same order as the video.

That's enough. You don't need to add "first the hand rises, then a turn, then the camera slowly orbits."

Two Kinds of Jobs: Anchored and Free

Seedance 2.5 quietly treats your job as one of two types, and knowing which one you're in saves a lot of confusion.

Anchored jobs — editing, first/last frame, and extending. Here your source material becomes an actual segment of the output timeline, so the result inherits its shape. Editing a clip keeps the original's aspect ratio and runs essentially the same length. Extending keeps the original's aspect ratio. Starting from a first frame keeps that image's aspect ratio.

Free jobs — reference-driven generation, storyboard grids, keyframes. Here the material is only a semantic reference, so you stay in control of the output's aspect ratio and duration.

Two things surprise people:

  • Storyboard grids are free jobs. A multi-panel storyboard image is treated as loose story guidance, not a strict frame-by-frame contract. The output will not match your panels exactly.
  • Keyframes are free jobs too, but they do align closely to your images. If you need tight visual fidelity to your drawings, use keyframes, not a storyboard grid.

One practical note on editing: if you want a clip trimmed, swapped, or repainted, your prompt needs to sound like an edit. Words like edit, add, remove, delete, change, replace are what signal the intent — "add some small animals to video 1," "change the person in video 1 to the person in image 1," "remove the background music from video 1."

Timestamps: Directing Second by Second

Seedance 2.5 responds to timestamps at one-second resolution. This is the biggest control upgrade over 2.0, which only understood shot numbers.

Three ways to use them:

Explicit ranges — keep the timeline continuous, with no gaps:

0-3s: ...
3-7s: ...
7-15s: ...

Or bracketed, if you prefer: [1s-4s] ... [4s-8s] ... [8s-12s]

Single moments — for a hit or a transition:

At 5s, a fast whip-pan to the left transitions the scene.

At 2s, a golden bolt of lightning tears down from the top of the frame.

Relative timing — when the beat depends on an action, not the clock:

Zhang San stands frozen in place; three seconds later, the people around him shake their heads.

After the protagonist presses the shutter, the frame holds for one second.

Budget your seconds honestly

Timestamps fail in two directions. Give a window too little content and the model improvises to fill it. Cram too much into a short window and you get rushed cutting or dropped beats. If a segment has three actions, a line of dialogue, and a camera move, it probably needs more than two seconds.

One thing timestamps are not for: frequency. Don't write "shakes their head three times in one second." Describe the action, not the count.

Saying No: Subtitles and Sound

Seedance 2.5 accepts negative instructions in two areas, and both are genuinely useful.

Subtitles:

No dialogue subtitles.

No subtitles.

Audio — and you can be selective here:

No background music; generate ambient and action sound only.

No sound at all.

That granularity matters. Plenty of ad and social work needs clean ambient audio with no music bed, and this is how you get it without stripping the audio afterward.

Camera Language That Works

Common terms — just write them. Shot sizes (extreme wide, wide, medium, close-up, extreme close-up), moves (push in, pull out, pan, track, follow, orbit, dive, tilt up, handheld shake), and angles (low angle, top-down, first person) are all understood directly.

Popular techniques — also fine as-is. One-take/oneer, dolly zoom, aerial view, FPV, bullet time, handheld, speed ramping.

Obscure or highly technical terms — explain them. Pair the name with a plain description of what happens on screen:

Rack focus: the focal point shifts smoothly — the tree that was sharp in the foreground goes soft, while the figure in the background comes into focus.

Transitions — name the trigger and the method:

At 5s, a fast leftward transition (left wipe with a natural dissolve).

Action and Expression

For action, summarize. Broad strokes beat exhaustive choreography: "does several sets of high knees and a backflip," "the two break into close-quarters combat." Save the detailed, frame-specific description for the one or two moments that need to land. Repeating the same action description over and over doesn't reinforce it — it muddies it.

For expression, describe rather than label. Idioms and shorthand don't translate well into faces. Instead of "eating with great relish," write:

A satisfied smile on their face, taking big mouthfuls of food.

How Much Material to Feed It

Fifty references is the ceiling, not a target. Practical limits and sweet spots:

MaterialCeilingSweet spot
Images30 images, up to 4K each
Video10 clips, 30s total
Audio10 clips, 30s total
Character refs from video/audio1-5 characters; 6-10 works but gets unstable
Character reference clip length5-10 seconds each
Character refs from images1-8 characters; 9-12 gets unstable
Storyboard panelsUnder 15 panels
Clip length for editingUnder 20 seconds
Reference images for an edit1-5 images; 6-8 gets unstable

A note on multi-view character sheets: for up to five characters, single-view or multi-view references both work. Past five, single-view is more stable — and if you do need multiple angles, upload them as separate images rather than one image containing a grid of views.

Task Playbooks

Keyframes — when the visuals must match your drawings

Upload your frames in order, and open the prompt by saying so explicitly:

Using images 1 through 7 in order as keyframes: across a sea of clouds and mountain peaks, a blue-and-pink long-tailed spirit fish glides through the air; the camera pushes slowly toward an ancient town built into the mountainside, settling on the pagoda at the summit; the scene cuts to an elegant hall, where the spirit fish flies in through a window and drifts into the round pool at the center; finally the view shifts to a dim old temple, where a white-bearded monk stands with his back to us, gazing at a huge framed painting — inside which is the hall and the fish in the pool. New-Chinese-style ukiyo-e illustration.

Output tracks the images closely. This is the tool for a storyboard you actually need honored.

Storyboard grids — loose story guidance

A multi-panel grid gives the model the shape of the story, not its exact frames. To use it well:

  1. Keep it under 15 panels. Eighteen-panel grids tend to produce frozen shots and scrambled order.
  2. Use simple line art or stick figures. Over-sharpened, cluttered, AI-generated storyboards hurt more than they help, and heavy text on the panels bleeds into the output.
  3. Fill in what the drawing can't say. Panels don't carry motion, camera, or style — write those in.

A clean structure: state your material mapping, give the overall story in a sentence, then walk the shots.

Material binding: storyboard @image1, bedroom @image2, Li Tian @image3,
Li Qian @image4, the book "Happy Days" @image5.

Shot 1: [Wide static shot, eye level, rule of thirds] A snowy winter night in a
room. At the floor-to-ceiling window, a man stands in profile with his hands in
his pockets, watching the snow fall. A young woman stands quietly beside him.
The mood is still and restrained; snowflakes keep landing on the glass.

Shot 2: [Medium over-the-shoulder] The young woman's back in the foreground. The
man turns his head and looks at her gently; she lowers her head in silence. Snow
continues outside.

Shot 3: [Medium close-up, diagonal composition] The man slowly holds out the book
"Happy Days"; she reaches up and takes it.

Shot 4: [Close-up on her face, centered] She hugs the book to her chest, eyes
reddening, a tear sliding down. Her expression is wistful.

Shot 5: [Close-up on his face, angled] The man smiles softly, watching her cry,
a trace of melancholy in his eyes.

Shot 6: [Wide static shot] She turns and walks slowly out of frame, leaving him
alone at the window, hands in pockets, looking out at the snow. The room is empty
and quiet.

If your panels are conceptual rather than literal, keep the prompt short and let the model work:

Following the order of the storyboard, build a complete story with coherent, natural camera work.

Blockouts — pre-staging your camera in 3D

You can hand Seedance 2.5 a rough grey-model video — untextured geometry — and have it render the final look while following your camera moves, timing, and motion paths. This is how you get deliberate framing instead of prompt-and-pray.

Two flavors:

Rough blockouts work best right now. Build with simple geometric shapes standing in for people, animals, and props. Keep the bodies simple — avoid limbs and wings unless you've animated their full motion, or you'll get stiff appendages in the render. State exactly what you want taken from the blockout:

Reference the camera work and motion from video 1.

Reference the lighting changes and camera work from video 1.

If you're also supplying character images, spell out which model each one maps to:

The man in the grey outfit in image 1 corresponds to the red model in video 1; the red-haired girl in video 2 replaces green model 2 in video 1.

And still describe the scene you want in full. A blockout carries structure, not content — the richer your written description, the better the render, as long as it matches what the geometry is doing.

Fine blockouts are for fully modeled scenes you want re-rendered — essentially "coloring in" a finished animation. Keep the source clean: no trajectory lines, no coordinate grids, no camera cones in frame, or they'll leak into the final image.

Editing — changing a video you already have

Be specific about the scope and the change, and describe the transformation from A to B. Timestamps let you confine the edit to part of the clip.

Only edit the man's dialogue in video 1 — change it to "Don't come any closer," with a northeastern accent.

Change the man drinking coffee at 4-6 seconds in video 1 to mopping the floor; leave everything else unchanged.

Edit task: replace the Asian woman on the right side of video 1 with the Black woman in image 1.

You can also feed reference images to guide an edit:

Replace the plain sparring footage in video 1 with an unarmed standoff before a cold-weapon duel. Replace the setting with a medieval stone keep terrace, backed by castle walls, wind, mist, and distant mountain lines, with a flat stone floor per image 1. Replace the dark-clothed man's outfit with image 2, and the light-clothed man's with image 3. Keep the choreography and original rhythm unchanged. Effects should only reinforce environment and texture: wind in the clothing, light mist, a little dust at contact points, cold metallic reflections, subtle grain, and epic color grading. Restrained, real, classical hard-edged duel atmosphere, with the music hitting on the beats.

And audio is editable on its own — voices, music, and effects can each be added, changed, or removed:

Translate the spoken dialogue in the video into Chinese, no subtitles, with lip movement precisely adjusted to match. Everything else stays the same.

Extending — continuing a clip forward or backward

Say which direction and how long, then describe the new material:

Continue video 1 with five more seconds: a bee flies in and lands on the flower, then a macro close-up of its legs and abdomen covered in golden pollen grains. The bee lifts off, the camera following it to another flower of the same kind. In slow motion, pollen shakes loose from its fuzz and lands precisely in the stigma — the moment of pollination, magnified.

Trigger words that signal an extension: extend forward/backward, continue, carry on. You can combine it with references too — "extend video 1 forward, with the character from image 1 dropping in from above."

One honest caveat: volume can shift slightly between the original and the extension. It's least noticeable when the source clip was generated by Seedance 2.5 itself, which is worth knowing if you're planning a long sequence.

Montage and seamless transitions

One-click montage turns a pile of stills or clips into a finished short:

Turn all the images into a single edit, arranging them freely, in a hand-drawn animated cutout-doodle style — a coffee shop vlog following a small dog in different cute outfits posing around the café. Generate playful, internet-native audio or background music. Let the images move subtly, like Live Photos, without altering the originals — stay closely faithful to each source image.

Seamless transitions bridge two clips by generating the connective tissue between them:

Join video 1 and video 2. Video 1's viewpoint flies rapidly up to the top and whips back, diving vertically downward, then transitions seamlessly and naturally into video 2. During the cut, the mahjong tiles gradually become high-rise buildings, and the whole scene changes to match. Do not alter either of the uploaded videos.

What Changed Since Seedance 2.0

If you've been prompting 2.0, four things are genuinely different:

  • Timestamps now work. Seedance 2.0 only responded to shot numbers; 2.5 responds to whole-second timestamps. Rewrite your beat structure to take advantage.
  • Multi-view character sheets are supported. 2.0 discouraged them; 2.5 handles them (with the caveat about five-plus characters above).
  • Aspect ratio is essentially free. 2.0 offered six fixed ratios. 2.5 supports anything from tall to wide, shaped by your input materials.
  • Editing and extending hold up better. Color, brightness, and audio-visual continuity are noticeably more consistent across an edit or a continuation.

Everything you already know about subject, motion, audio, and style references carries straight over. The new capabilities sit on top of the old skills, not in place of them.

Common Mistakes Checklist

Run through this before you generate:

  • Mapping is in the image, not the text. Write every reference role in the prompt.
  • Broken timeline. "0-3s ... 5-6s" leaves a hole. Keep ranges continuous.
  • Overstuffed segments. Five beats in two seconds means dropped beats.
  • Describing what the reference already shows. Competing instructions produce mush.
  • Contradictions between the top and bottom of the prompt. Long prompts drift — reread the closing paragraph against the opening.
  • Too many storyboard panels. Over fifteen and things freeze or scramble.
  • Repeating the same action detail. Say it once, clearly.
  • Idiomatic expression descriptions. Describe the face, not the phrase.
  • Using timestamps to control frequency. Describe the action instead.

Conclusion

Seedance 2.5's real leap isn't just that a clip can run thirty seconds. It's that thirty seconds is long enough to hold an actual story — and the model now gives you the controls to direct one: second-by-second timing, dozens of mapped references, precise editing, and clean continuation.

Those controls only respond to prompts written like direction. Bind your materials. Summarize the piece in a line. Lay out the timeline. Close with what must hold throughout. Then go specific where it counts and stay out of the model's way where your references already speak.

Start simple: take a prompt you've written before, add a material mapping block and a timeline, and generate it again. The difference is usually obvious on the first try.

Ready to direct? Seedance 2.5 is available now on HeyMarmot — pick it from the model list, upload your references, and start writing your shot list.