Choose the starting composition
Use an image with a clear subject and enough context for the intended movement. A tight crop may not provide the same room for action as a wider view.
Video / Use case
Use a still image as a visual reference and focus the video brief on what should move, what should stay, and how the shot resolves.
An image can already express the subject, framing, and mood. Instead of rewriting everything visible in it, use the instruction to explain the intended movement. Keep the action simple enough to judge the result against a clear intention.
From input to intention
Prepare the input, express the intention, and review what actually changed.
Use an image with a clear subject and enough context for the intended movement. A tight crop may not provide the same room for action as a wider view.
Describe the subject’s action, any camera movement, and the elements that should stay stable. Keep the requested motion compatible with the visible scene.
Compare the visual identity, structure, and mood with the reference. Watch for unwanted changes through motion rather than judging only the opening frame.
Start with a useful brief
Use this as a starting brief, then adjust it to your own source and intention.
integrated_multimodal_description: [Shot 1] Start exactly from the supplied image: a clear perfume bottle with a black cap and pale golden liquid stands on a dark stone plinth beside a softly lit window. A slow, restrained camera push approaches the bottle over five seconds. The bottle and plinth remain still. Glass reflections shift subtly and naturally with the camera. Preserve the bottle shape, cap, liquid and room throughout the shot. No cuts, hands, extra objects, label, lettering, logos or watermark. overall_soundscape: Quiet indoor room tone. No speech or sound effects. non_diegetic_music: N/A
Exact recorded Image to Video prompt. The recipe includes the starting product still, five-second duration, square framing and detailed rendering controls.
A workflow may use an image as guidance without guaranteeing that the first output frame is an exact copy. A dedicated start-frame control expresses a more specific requirement when the backend supports it.
Do not label every image input “first frame” unless that is how the connected mode actually works. The public description should match the implemented input semantics.
A still portrait can suggest small expressions or turns, while a wide landscape may suit camera movement or subtle environmental motion. The source should give the intended action enough context.
If the desired movement reveals areas not visible in the image, inspect how those areas are generated. A convincing starting view does not guarantee consistent unseen structure.
Use Video Extender when a successful clip needs a continuation. Use Video Edit when the existing clip needs a visual revision. Return to the source image when the composition itself is the problem.
Talking Face is a more specific reference workflow involving photo and audio. It should not be implied by every image-led video generation.
The useful details
Not necessarily. That depends on whether the connected mode treats it as a reference or supports a dedicated starting frame.
Under Video Generator. In the Video Generator family, it is the Images to Video mode.
It is a specific photo-and-audio use case inside Video Generator, with additional performance and permission considerations.
Focus on motion, pacing, camera behavior, and what should remain stable rather than repeating every visible detail.
Connected workflows
One tool. More possibilities.
This use case belongs inside the broader tool. Explore related approaches without learning a different navigation system.