Seedance 1.5 Pro is a video model that produces visuals and clean audio at the same time. In this guide we'll break down how to write prompts for Seedance so you can use its strengths:
- Native audiovisual synchronization;
- The ability to perform complex camera moves and auto-plan motion and transitions;
- Multi-shot character consistency.
You'll learn how to avoid conflicts between the visual track and the sound, which details to put first for maximum model attention, and get ready-made prompts for creating videos. We'll show why Seedance 1.5 needs a more structured approach to prompting than previous models — and why it pays off.
New Seedance 1.5 Pro features and how to use them in prompts
Every new model feature changes how you should phrase a request — ignoring these capabilities leaves half the tool's potential on the table.
Audiovisual synchronization
The main advantage of Seedance 1.5 Pro is its ability to predict sound from on-screen motion without explicit instructions. The AI analyzes object dynamics and generates a matching audio track: fingers moving across guitar strings produce specific notes, a car screeches tires exactly at the moment it skids, breaking glass sounds in sync with the shards.
The model understands the link between visual dynamics and sound patterns thanks to training on large audio-video datasets. Sync quality is one of the model's headline strengths, which is why contradictions like "a quiet storm with powerful lightning" create a conflict — the model can't decide whether to prioritize "quiet" or "powerful."
Camera control
To control camera movement, describe it directly in plain words in your prompt: a pan left to right or right to left, a tilt up or down, a smooth push-in or pull-out, a slight roll around the axis (as if the camera is "tilting the horizon"). Spell out movement, speed, and direction literally: "smooth pan left to right," "slow tilt up toward the sky," "gradual push-in on the face," "slight clockwise roll of a few degrees."
For a dolly zoom effect, describe two opposite movements at once: "the camera slowly pulls back, but the character's face stays roughly the same size in frame — the background visually 'stretches,' creating an unsettling effect."
For a tracking shot, specify the trajectory and position: "the camera follows a running dog left to right at eye level, keeping it centered in frame."
For an orbital shot, set the arc and angle: "the camera smoothly orbits a standing object 180° counterclockwise, keeping the object centered in the composition."
Here's a sample prompt:
"An evening city street, soft light from shop windows, wet asphalt after rain. One adult (over 25) stands at the curb, looking ahead. Smooth pan left to right: the camera slowly pans left to right, showing the street and returning to the person, who remains the main subject. Tracking shot: the person starts walking, and the camera follows alongside left to right at eye level, keeping them centered in frame. Dolly zoom (brief): at the end — a dolly zoom effect: the camera pulls back slightly, but the person's face stays roughly the same size in frame while the background visually stretches a little. Sound: quiet city ambience (distant cars, footsteps, a light breeze). No music, no speech. All camera movements are smooth, with no jerks or shake."
Customize the prompt in Creative Lab →
Combine movements for complex cinematic effects: "the camera slowly rises while simultaneously pulling back, revealing the cityscape behind the character in the foreground." Combined prompts like this use the full potential of the camera control system.
Lip sync for talking characters
Seedance 1.5 supports English, Chinese, Japanese, Korean, Spanish, Indonesian, and Portuguese. It can distinguish tone — an important feature for Asian languages, where pitch changes word meaning. For a talking character, specify the language, pace, and line in your text prompt. For example:
- "A man speaks English slowly and clearly, looking into the camera: 'Hi everyone. Today I'll show you how this works.'"
- "A young woman speaks Spanish quickly with emotional intonation, gesturing with her hands: '¡No puede ser! ¡Es increíble! ¡Probémoslo ahora mismo!'"
Avoid unnatural camera angles when lip syncing. For example, "a frontal close-up of the face with soft lighting" works more reliably than "a side profile in deep shadow," because the model was trained mostly on frontal talking-head footage.
Character appearance
Seedance 1.5 Pro lets you upload a reference image of a character, object, or location. This lets the model "remember" what the scene looks like and preserve its key details from shot to shot, even if the setting or angle changes. That helps you create consistent videos with the same characters without manually matching every scene.
For example, you can generate a character image using this prompt:
"Cinematic shot, a young European man around 30, gray hoodie and round black glasses, short hair, light stubble, standing relaxed with hands in pockets, looking straight into the camera, dramatic background with an ocean cliff at sunset, waves crashing below, golden-orange sky with purple clouds, wind gently moving his hair, medium shot from the waist up, epic cinematic lighting, rich warm tones, photorealistic."
After that, attach the image when creating the video and reference the same character: "The man from the reference stands on the cliff at sunset. Wind blows through his hair. He squints and smiles. Sound effects: wind, ocean waves." The model will preserve the facial features, glasses, and clothing style, adapting only the pose and setting. Here's the result of the prompt:
For reliability, avoid drastic lighting changes between scenes: a transition from a bright café with large windows to a night street with neon light should be smooth, passing through an intermediate dusk scene. A character created in a daytime coffee shop will keep their facial features and clothing details when moving to an evening city street, but may lose their appearance in an abrupt jump to a night scene with colored lighting.
Principles of effective prompting for Seedance 1.5 Pro
Prompt structure determines the result more than the creativity of the wording — the model expects a certain order of information and prioritizes elements by their position in the text.
Six essential prompt components
Any prompt for Seedance, on the free or paid tier, should include the blocks below, arranged in a strict sequence for the model's optimal attention.
- Subject: a specific character or object with a detailed description — "a woman astronaut in a white spacesuit" instead of a vague "a person" or "someone."
- Action: what the subject is physically doing — "walking slowly," "stops and touches her helmet," "looks back over her shoulder."
- Facial expression: a detailed description of the expression to convey emotion — "her brows furrow, her eyes dart nervously, her breathing quickens" instead of a simple "looks worried."
- Dialogue (optional): the character's line with a tone of voice — "[Woman astronaut, trembling, anxious voice]: 'I think I just lost my signal.'"
- Camera movement: how the camera follows the action — "the camera follows alongside at eye level, slowly moving in" or "the camera holds a close-up on her face."
- Atmosphere and sound: visual style plus sound effects — "cold blue light with harsh shadows, a sci-fi atmosphere. Sound effects: heavy breathing inside the helmet, crackling radio static, a warning beep from the suit."
These elements form a minimum viable prompt that will give a predictable result.
Example of a full prompt: "A woman astronaut in a white spacesuit walks slowly across a red Martian desert. She looks back over her shoulder with a worried expression. Her brows furrow. Her eyes dart nervously. Her breathing quickens inside the helmet. She stops and touches her helmet with one hand. She says: 'I think I just lost my signal.' The camera follows alongside at eye level, slowly moving in. Cold blue light with harsh shadows from the red rocks, a sci-fi atmosphere. Sound effects: heavy breathing inside the helmet, boots crunching on sand, a distant wind howling, crackling radio static, a warning beep from the suit, an accelerating heartbeat."
This structure gives the model a clear hierarchy: focus on the astronaut, her emotions drive the pacing, the camera emphasizes isolation, and the lighting sets the genre.
How to avoid conflicts between the visual and audio layers
The model links visual motion to sound through learned patterns: fast movement automatically generates dynamic sound, slow movement generates calm ambience. Contradictions in a prompt create uncertainty that Seedance 1.5 Pro resolves unpredictably. Avoid phrases like "a quiet storm with strong wind and bright lightning" — that's three conflicting signals:
- "Quiet" implies faint sound;
- "Strong wind" calls for a loud whistle;
- "Bright lightning" expects rolling thunder.
Instead, pick a dominant trait: either "a distant storm with quiet thunder rumbles and a light breeze," or "a powerful storm with loud thunderclaps and roaring wind."
To control sound, name its source explicitly: "a guitarist plays a slow acoustic melody on a classical guitar" instead of a vague "music plays in the background." If you need silence or minimal sound, describe a static scene: "a person stands motionless in an empty white room, the only sound is the quiet ticking of a wall clock."
When generating video with audio, the visuals should match the track's pace and mood: fast speech calls for active gesturing, a slow lecture calls for restrained movement.
Working within the model's limits: duration and complexity
In GPTunneL, Seedance 1.5 Pro generates video in two formats: 5 or 10 seconds. You can choose the video length, aspect ratio, quality (720p or 480p), and whether to add audio.

If you need a longer scene, split it into a sequence of 10-second clips and tie them together with shared logic: repeat the description of the same character, attach a reference for their appearance, clothing, setting, and key actions to keep continuity.
Ready-made prompt examples
Ready-made prompts help you quickly pick up the model's syntax and understand how it interprets different types of tasks — from vertical social content to cinematic previsualization.
A street musician and city sounds
Context: A content creator makes an atmospheric short scene where action sounds matter: strings, coins, a light breeze. The goal is to show how the model links visual events to specific sound effects.
Prompt: "A street musician sits by a brick wall during the day, playing an acoustic guitar. The camera slowly pans left to right, moving from his hands on the strings to his face. As a passerby drops a coin into the open case, a distinct metallic clink is heard. Sounds: rhythmic acoustic guitar, light city ambience, footsteps on the pavement, occasional distant voices, a brief gust of wind. Soft, natural light, documentary aesthetic."
Result: A 10-second vignette where the sound of the strings matches the hand movement, the coin's "clink" lines up with the drop, and the city ambience feels believable. Customize the prompt in Creative Lab →
A night forest after rain
Context: The author creates an atmospheric forest video with no people in it. The sound needs to be realistic and synchronized: wind moves the leaves, there's rustling, a fox passes by, and a stream flows nearby.
Prompt: "A night forest after rain. Close-up of wet leaves and branches, water droplets falling. The camera slowly tilts from top to bottom, revealing a small stream flowing over rocks. In the mid-ground, in the shadows behind the bushes, a fox passes quickly and quietly. Branches and leaves stir slightly as it passes. Sound: a quiet rustle of leaves exactly as the fox passes, light footsteps on wet ground, a constant stream sound, occasional 'tick' drops, one distant owl call. Moonlight, light fog, realistic."
Result: 10 seconds of forest atmosphere with realistic water and foliage sound, a brief dynamic moment with the fox, with the rustling and footsteps matching the on-screen motion. Customize the prompt in Creative Lab →
A woman restoring a mosaic in a restoration studio
Context: A short-video creator is making a clip for a museum or gallery. It needs to demonstrate Seedance's strengths: synced sounds (stone, glue, tool), accurate object rendering, a calm line with English lip sync, and controlled camera work.
Prompt: "A quiet restoration studio. An ancient mosaic made of small stones lies on the table, with small tools and a brush nearby. In frame is an adult woman, 35–45, calm, focused, without theatrical drama. Action: she carefully picks up one small stone with tweezers, places it into an empty spot in the mosaic, and presses it gently. Then she looks up at the camera and speaks in English, quietly, quickly and clearly (once, no repeats): 'When you see the pattern, every piece knows its place.' After the line, she looks back down at the mosaic.
Camera: first a close-up of her hands and the stone, then a smooth transition to her face for 2 seconds during the speech, then back to the mosaic. No sudden movements. Sound: a quiet scrape of tweezers, a light 'click' of the stone settling into place, the soft rustle of the brush on the surface, a calm room tone. The voice is clear and up front. No music, no impacts, no explosive effects. Rule: the line is spoken once, without duplication."
Result: A scene where the sound precisely matches the stone being set, the voice sounds natural and in sync, and the editing creates a clear sequence of shots: first the hands, then the character's face, and finally the mosaic itself. Customize the prompt in Creative Lab →
A talking cat with lip sync
Context: A user is making a short meme clip for TikTok to show off Seedance's strengths: a clear English line with "lip" sync, expressive vocal delivery, and atmospheric ambient sound (torches, hall), while keeping the scene simple to generate.
Prompt: "Premium cinematic realism. A large throne room in an old castle, warm torchlight, tall stone walls, a rug on the floor. In the center — a cat sits on a tall throne. The cat wears a small crown and looks confident and 'regal.' Camera: medium shot, then a very slow push-in on the cat's face. Action: the cat sits still and majestic, slightly raises its chin, and moves its mouth as if speaking. It says one short line in English in a commanding, authoritative adult voice (once, no repeats): 'Bring me my treats. Now.' Sound: a low reverberant ambience of the throne room, quiet crackling torches, a light rustle of fabric on the throne. The cat's voice is clear and up front. No music, no extra words."
Result: The cat on the throne looks regal, the camera smoothly pushes the viewer in on its face, the line is spoken once, and the ambient sounds (torches, hall) support the atmosphere without extra noise or repeated lines. Customize the prompt in GPTunneL →
Conclusion
Seedance 1.5 Pro requires a balance between detail and clarity: describe the subject, action, camera movement, and lighting style in a structured way, but avoid overloading the prompt with unnecessary adjectives, complex instructions, or extra lines. Start with the basic four-part formula (subject + action + camera + style), then experiment with cross-modal details: how visual action affects sound, how speech pace changes a character's gestures.
Once you start working with Seedance 1.5 now available in Creative Lab, over time you'll learn how to create real masterpieces with this model.
FAQ
How do I keep a character consistent across multiple scenes?
Upload a reference image so the model knows what the character should look like. Then in following prompts use the phrase "the character from the reference" with a brief reminder of key details. Avoid drastic lighting changes between scenes — smooth transitions across times of day help maintain visual consistency.
Which languages can generate voice in a video?
The model supports English, Chinese, Japanese, Korean, Spanish, Indonesian, and Portuguese.
What should I do if a prompt isn't working as expected?
Simplify the prompt by removing elements one at a time: first drop the camera movement instructions, then remove secondary visual details, then simplify the audio component. Check for contradictions between the pace of the action and the sound. Rephrase abstract concepts into concrete visual scenes with clear objects and actions.
