VEO3 Prompt Guide: 8-Second Video with Sound

VEO3 is the new version of the video generation model from Google DeepMind that lets you create realistic 8-second clips with synchronised sound, including dialogue, sound effects and background music. All of it is generated from a single prompt, with no programming or editing skills needed.

An overview of what VEO3 can do

VEO3 offers a wide range of features for creating video content:

  • Video generation: Creating 720p video from text prompts
  • Audio generation: Automatically adding and syncing dialogue, voice-over, music and sound effects with the footage
  • Understanding complex prompts: Better interpretation of details, nuances and long instructions
  • Cinematic control: The ability to control camera parameters (panning, zoom, angles), lighting, cinematic styles and shot composition
  • Consistency: Keeping characters and objects visually and stylistically consistent across several scenes
  • High-quality lip sync
  • Realistic facial expressions and emotions
  • Support for non-English dialogue
  • The ability to work around some content restrictions

Principles of writing prompts

3.1 Core principles

Detail: Give descriptions that are as detailed as possible. Include details about the scenes, the characters (appearance, clothing, emotions), the actions, the surroundings, the lighting (for example, "golden hour", "gloomy lighting"), the colour palette and the overall atmosphere.

Structure and clarity: State the result you want clearly. For complex scenes you can break the description into logical parts or sequential instructions.

Language: We recommend writing prompts and audio inserts in English, since the model was originally trained on large volumes of English-language data. To generate speech in a specific language, you can use clarifying constructions in the prompt.

Sound specification: Explicitly name the sound effects, the music (genre, mood, instruments) you want, or state that you need silence. For dialogue, write out the characters' lines clearly.

3.2 Basic prompt structure

Length of the generated video: 8 seconds

Prompt formula:

[Character] + [Appearance] + [Action] + [Speech/Sound] + [Setting] + [Style/Quality]

4. A detailed breakdown of the components

4.1 Character + Appearance

The central character of the video with a detailed description of appearance, age, clothing and distinctive features. This is the foundation of the generation: it defines who will be in focus and how they look.

Examples:

  • "A thirty-year-old man with dark beard wearing vintage military uniform"
  • "Young woman with long blonde hair in red evening dress"
  • "Elderly professor with glasses and tweed jacket"
  • "Teenager in bright pink hoodie with curly hair"
  • "Middle-aged woman with short black hair in business suit"

4.2 Action

The character's key physical action or behaviour

⚠️ CRITICAL: if you plan to have speech, the character can ONLY speak (says/sings) — no other actions at the same time!

For speech (one action only):

"says", "sings", "whispers", "shouts", "speaks"

For actions without speech:

  • "walks slowly", "dances gracefully", "sits contemplatively"
  • "gestures dramatically", "looks around nervously", "smiles warmly"
  • "moves chess piece", "opens door carefully", "writes in notebook"

4.3 Speech/Sound

Dialogue, sound effects and the audio atmosphere. Non-English speech must always go in quotation marks after you name the language.

Formats for non-English speech:

  • says in Spanish: «Hola, ¿cómo estás?»
  • sings in Italian: «Volare, oh oh, cantare, oh oh oh»
  • speaks with a southern accent in Portuguese: «Amigo, está tudo bem»

Sound effects:

  • "with ticking clock sounds in background"
  • "ambient forest sounds with birds chirping"
  • "dramatic orchestral music swelling"
  • "footsteps echoing in empty corridor"
  • "gentle rain sounds and thunder"

4.4 Setting

The location, surroundings and context where the action unfolds. It creates the spatial and temporal context for the character.

Examples:

  • "in dark gothic room with wooden furniture and candlelight"
  • "standing in bright modern office with large windows"
  • "sitting in cozy cafe with warm lighting and vintage decor"
  • "walking through snowy forest path with tall pine trees"
  • "in underground bunker with concrete walls and dim lighting"

4.5 Style/Quality

Technical quality, cinematic techniques, lighting, colour palette and the overall aesthetic of the video.

Technical parameters:

  • "cinematic quality, high resolution"
  • "medium shot", "close-up", "wide angle"
  • "handheld camera movement", "steady cam"

Camera control:

  • "close-up shot of a character's face" (a tight shot of the character's face)
  • "wide aerial shot of a landscape" (a wide panoramic aerial shot of scenery)
  • "drone shot following a car" (a drone shot trailing a car)
  • "slow panning shot across the room" (a slow pan across the room)

Lighting and style:

  • "dramatic shadows with noir-style lighting"
  • "warm golden hour lighting with soft shadows"
  • "cold blue moonlight creating mysterious atmosphere"
  • "vintage film grain with muted colors"
  • "sharp contrast between light and dark areas"

Stylisation:

  • "cinematic lighting" (cinematic lighting)
  • "vintage film look" (the look of old film stock)
  • "watercolor style" (watercolour style)
  • "photorealistic" (photorealistic)

Mood:

  • "melancholic autumn atmosphere"
  • "energetic and vibrant mood"
  • "suspenseful thriller aesthetic"
  • "romantic soft-focus ambiance"
  • "dystopian cyberpunk atmosphere"

5. Examples of finished prompts

5.1 Classic examples with non-English dialogue

A cosmonaut:

A fifty-year-old veteran cosmonaut with grey hair wearing white spacesuit with colourful mission patches, floating in space station cockpit with blinking control panels and says in Spanish: «Tierra, veo la Tierra, es hermosa», with wonder and awe in his voice, surrounded by glowing instrument displays and Earth visible through porthole window, ambient space station humming sounds, retro sci-fi aesthetic with warm orange lighting, cinematic quality, medium shot, high resolution, nostalgic space-age atmosphere.

A chef:

A thirty-five-year-old Mediterranean chef with thick black mustache wearing white chef hat and apron, standing in busy restaurant kitchen with steaming pots and says in Italian: «Amici, la focaccia è pronta, assaggiate», with proud smile and welcoming gesture, surrounded by sizzling pans and chopping sounds, warm kitchen lighting with copper pots gleaming, traditional restaurant ambiance with clinking dishes background, medium close-up shot, high resolution, cozy culinary atmosphere.

A ballerina:

A twenty-five-year-old prima ballerina with elegant bun hairstyle wearing white classical tutu and pointe shoes, standing center stage of grand opera theater and says in French: «Ce soir je danse pour vous de tout mon cœur», with graceful posture and emotional intensity, surrounded by red velvet seats and golden baroque decorations, soft orchestral music swelling in background, dramatic stage lighting with warm spotlights, cinematic quality, wide shot transitioning to medium, high resolution, romantic theatrical atmosphere.

5.2 Examples for different goals

1. Creating a short scene with dialogue:

A dimly lit, old library. An elderly historian with glasses perched on his nose looks up from a large, ancient book and says with a thoughtful expression: 'The secrets of the past are often hidden in plain sight.' Soft rustling paper sounds in the background

2. Generating with a specified style and camera movement:

A hyper-realistic drone fly-through of a lush, alien jungle at twilight. Bioluminescent plants glow faintly. Eerie, atmospheric alien sounds

3. Creating an atmospheric scene with the emphasis on sound:

Time-lapse of clouds moving across a stormy sky over a rugged mountain range. The wind howls, and distant thunder rumbles. No music

4. A product promo clip for marketing:

A sleek smartphone rotates slowly on a white pedestal, studio lighting creating dramatic reflections on its glass surface. Camera performs a 360-degree orbit. A confident female voice says: 'Innovation meets elegance.' Subtle tech ambient music builds. The phone screen illuminates showing colorful app icons

5. An educational science visualisation:

Microscopic view diving into a human cell. Camera zooms through the cell membrane, past floating organelles. Narrator with British accent explains: 'The mitochondria, often called the powerhouse of the cell, produces energy through cellular respiration.' Soft electronic music, subtle bubble sounds

6. An advert for social media:

Fast-paced montage: barista pours latte art in slow motion, steam rises dramatically. Cut to: customer's eyes widen with delight. Text overlay appears: 'Morning Magic'. Upbeat acoustic guitar, coffee shop ambiance. 15-second format, vertical aspect ratio

7. A corporate presentation:

Modern glass office building exterior, sunrise time-lapse. Transition to: diverse team collaborating around holographic display. Professional woman in business suit turns to camera, smiles warmly: 'At TechCorp, we're building tomorrow's solutions today.' Corporate ambient music, subtle keyboard clicks

5.3 More non-English examples

A northern hunter:

A forty-year-old northern hunter with thick beard wearing fur hat and leather jacket, sitting by crackling campfire in snowy boreal forest and says in German: «Morgen bei Sonnenaufgang gehen wir auf die Jagd, seid bereit», with serious determination in weathered face, surrounded by tall snow-covered pine trees and dancing flames, wind howling through branches and wood crackling sounds, cold blue moonlight contrasting warm firelight, cinematic quality, medium shot, high resolution, harsh wilderness atmosphere

A Parisian intellectual:

A sixty-year-old Parisian intellectual with silver beard wearing vintage glasses and dark wool coat, walking along a grand boulevard in autumn evening and says in French: «Cette ville a toujours inspiré les grands écrivains», with contemplative wisdom in his voice, surrounded by classical architecture and golden street lamps, footsteps echoing on wet cobblestones and distant church bells, warm amber lighting with soft shadows, film noir style, tracking shot following character, high resolution, melancholic literary atmosphere

A steppe chieftain:

A forty-five-year-old steppe chieftain with long mustache wearing traditional fur hat and red kaftan with golden braids, mounted on black horse in vast steppe landscape and says in Turkish: «Halkımız ve toprağımız için, ileri süvariler», with commanding authority and pride, surrounded by endless grasslands under dramatic storm clouds, horse snorting and wind whistling across plains, epic orchestral music building, dramatic lighting with sun breaking through clouds, cinematic quality, low angle shot, high resolution, heroic historical atmosphere

6. Working with dialogue

6.1 Recommendations for dialogue

For monologues (recommended):

  • One main character
  • One piece of speech
  • Fewer potential glitches

For dialogue:

  • Split direct speech between characters clearly
  • Tie every line to a specific character
  • Use a gender split (man and woman)
  • Avoid phrasing like "the first one says, the second one says"

An example of correct formatting: A photographer holding a camera says: "How are you?"

7. Technical limits and tips

7.1 Character limits in direct speech

❌ Do not use:

  • Three exclamation marks (!!!)
  • Ellipsis (...)
  • Em dash (—)

✅ Use:

  • The exclamation mark (!)
  • The question mark (?)
  • The comma (,)
  • The full stop (.)

7.2 The balance of languages in a prompt

  • Non-Latin script should make up no more than 15-20% of the total prompt
  • If you get the error "VEO3 does not support this language", expand the prompt with details in Latin script

7.3 Frequent problems

  • Phantom subtitles — they can show up even when you specify "no subtitles"
  • Voices blending in dialogue between characters of the same gender
  • The dash gets read out as a separate word

8. Common mistakes

8.1 The main mistakes to avoid

Not enough detail in the prompt: Requests that are too general or vague can lead to unpredictable, irrelevant or generic results. The more precise and detailed the prompt, the better the model will grasp what you have in mind.

Contradictory instructions: Avoid mutually exclusive descriptions or requirements within a single prompt. For instance, do not ask for a "bright sunny day" and a "gloomy, foggy atmosphere" at the same time.

Ignoring the audio capabilities: VEO3 has powerful sound generation capabilities. Do not forget to state what you want from the sound design (music, effects, speech) when it plays an important role in the scene.

Overloading with action while generating speech: If the main focus is on the character's dialogue, try not to overload the prompt with lots of other complex simultaneous physical actions for that same character, since it can affect how natural the speech sounds and how well it syncs.

9. Capabilities and limits

9.1 Capabilities

  • High-quality lip sync
  • Realistic facial expressions and emotions
  • Background music and sound effects
  • Support for non-English dialogue
  • The ability to work around some content restrictions
  • Video generation at 720p
  • Cinematic camera control
  • Visual and stylistic consistency

9.2 Limits

  • Video length: 8 seconds
  • Occasional errors with subtitles
  • Restrictions on explicit content (though softer than competitors')

10. Conclusion

Mastering prompt engineering for VEO3 is an iterative process that takes practice and experimentation. Use the recommendations here as a starting point for writing detailed, clear and precise requests.

Remember the key principles:

  • Maximum detail in your descriptions
  • A clear prompt structure
  • The right balance between English and the target language
  • Precise instructions for the sound design
  • Avoiding contradictory instructions

And if you run into trouble writing prompts for VEO3, use our assistant to build the right requests. Just describe what you want to get, and it will put together a detailed prompt with correctly formatted dialogue and sound effects!

Try it in GPTunneL