Choose the role of each input
Decide which reference supplies appearance, motion, delivery, or timing. Use only combinations supported by the connected mode for the request.
Video / Use case
Explore image, audio, and video references within one mode of Video Generator, including Talking Face.
References are different ways to communicate a creative intention. An image can anchor appearance, audio can provide a performance cue, and a video can express movement or timing. Give every input a clear role so the direction stays coherent.

Connected workflows
From input to intention
Prepare the input, express the intention, and review what actually changed.
Decide which reference supplies appearance, motion, delivery, or timing. Use only combinations supported by the connected mode for the request.
Explain how the references should work together and what should remain stable. Resolve conflicts before asking the tool to interpret several competing signals.
Compare the result with each input’s intended role. Inspect identity, movement, timing, and scene relationships rather than scoring every output by visual style alone.
Start with a useful brief
Use this as a starting brief, then adjust it to your own source and intention.
Use the image reference for the subject’s appearance and the supplied motion reference for broad movement. Keep the environment visually restrained so the action is easy to read. Review the result for consistent identity and movement before continuing.
Illustrative creative brief. It is not the recorded prompt for the displayed sample.
Images to Video, Videos to Video, Audio to Video, and Talking Face remain discoverable as detailed pages. They belong to Reference to Video inside Video Generator, not as a row of new sidebar tools.
The pages explain differences in preparation and review. After integration, each can open the same tool with the relevant mode selected, without changing the broad navigation.
A picture and a video can disagree about the subject’s appearance or environment. Audio may imply a performance that does not fit the visual reference.
State which input should control each part of the intention. When the controls cannot express that separation, simplify the inputs instead of assuming the model will resolve the conflict in the intended way.
Explore the full family of reference workflows here and open the matching Studio entry. The connected service confirms supported input combinations and the quote before generation.
Every reference workflow has a detailed page and a matching Studio entry. Choose the inputs that communicate your intended appearance, motion, or performance.
Choose the right approach
| Your intention | A useful starting point | What to watch |
|---|---|---|
| Establish the subject and composition | Images to Video | A reference does not automatically guarantee the exact opening frame. |
| Describe a performance or a camera move | Videos to Video | Decide whether the source is an appearance reference, a motion reference, or the footage being edited. |
| Provide speech, rhythm, or timing | Audio to Video | Give the sound a visual intention rather than assuming that audio alone defines a face or scene. |
| Combine an authorized face and spoken performance | Talking Face | Review facial movement, synchronization, expression, and permission for both the face and the voice. |
| Connect two intended visual states | Start & End Frames | Review the transition as well as the two boundaries. Confirm exact frame support in the connected workflow. |
The useful details
No. Reference to Video is a mode inside Video Generator, with more specific entry pages for different reference intentions.
Video Generator → Reference to Video → Talking Face. It uses the photo-and-audio entry point rather than a new sidebar tool.
Do not assume that. The final connected backend must determine the supported combinations and input requirements.
Open this workflow in Studio to prepare its inputs, direction, and project assets. Review the supported options and the service quote before generating.
Connected workflows
One tool. More possibilities.
This use case belongs inside the broader tool. Explore related approaches without learning a different navigation system.