24/09/2026
Your reference image already describes the scene. Your video prompt should direct what changes. (Video in 1st comment)
When turning a still into video, repeating every visible detail can bury the actual motion instructions. A more useful structure is:
1. Subject action
2. Camera movement
3. Environmental movement
4. Audio in a separate sentence
Example:
“Eye-level medium shot. The bartender slowly turns toward the camera and places the glass on the counter. The camera performs a gentle dolly-in. The curtain moves slightly and reflections travel across the polished bar. Audio: quiet room tone and a soft glass-on-wood tap.”
The reference image carries the character, setting and composition. The prompt gives the shot direction.
Start with one clear action and one camera movement. Add secondary motion only where it supports the scene.
This doesn’t guarantee perfect consistency: the model may still alter faces, hands or object geometry. Google also notes that natural, consistent spoken audio—especially short speech—remains an area of active development.