Four sound layers you can describe
- Dialogue
- Write the exact words, the speaker and delivery. Keep lines short enough for the clip duration.
- Ambience
- Name the place and distance: close rain on a roof, distant traffic, room tone or a quiet crowd.
- Sound effects
- Attach effects to visible events, such as a splash at impact or a latch clicking when a door closes.
- Music
- Describe mood, instrumentation and whether music sits under dialogue; avoid requesting copyrighted songs or artist imitation.
An audio-directed prompt
A cyclist turns into a rain-soaked street at night. The camera tracks low beside the wheels, then slows before the rider exits frame. Cinematic realism. Sound: tires cutting through water in the foreground, light rain on metal awnings, distant traffic and a low electronic pulse without vocals.Use this prompt →
The sound clause identifies foreground, background and an explicit “without vocals” constraint. That gives the model a clearer mix than a generic request for cinematic sound.
Synchronization limits
Native audio can align effects and ambience with the generated action, but exact dialogue performance, lip movement and multi-speaker timing remain probabilistic. Short lines and visible speakers are easier to evaluate. Always review generated audio before publishing, especially for brand, safety or factual claims.
- Keep dialogue short and identify the speaker.
- Place effects on visible actions instead of vague moments.
- State important exclusions such as no music or no vocals.
- Regenerate or edit externally when frame-accurate timing is required.
