A useful music-video prompt gives the system decisions it can visualize while leaving the song to control timing. Clear nouns and observable actions matter more than long lists of adjectives.
Use a six-part prompt structure
Build the prompt in this order: setting, subject, action, lighting, camera behavior, and texture. Add one sentence for continuity or exclusions only when necessary.
Example: “An anonymous dancer crosses a mirrored salt flat at sunrise. Long shadows, pale coral light, slow tracking camera, fine 35mm grain. Keep the same silver coat in every scene.”
- Setting: where the scene exists.
- Subject: who or what viewers follow.
- Action: what visibly changes.
- Lighting: time, direction, and contrast.
- Camera: framing and movement.
- Texture: film, material, or finish.
Match language to musical energy
For quiet passages, describe controlled camera movement, restrained action, and longer visual beats. For high-energy sections, use kinetic actions, sharper camera movement, and stronger changes in scale.
Avoid telling the model to “sync perfectly.” The generation workflow already uses musical structure; your prompt should focus on what changes when the energy changes.
Protect continuity with a few invariants
If a character appears across scenes, define only the traits necessary to recognize them: silhouette, clothing, hair, and one distinctive color. Repeating a short stable description is more effective than adding new details in every scene.
Use negative constraints for likely failure modes such as extra people, visible text, logos, abrupt wardrobe changes, or facial close-ups.
Refine one variable at a time
When the result is close, change one part of the direction instead of rewriting everything. Adjust camera speed, setting, palette, or shot scale individually so you know which change improved the video.
Save successful directions as reusable templates. A small prompt library becomes a practical creative system for recurring channels and campaigns.