Video / Use case

A face. A voice.
A moment of expression.

Plan a photo-and-audio reference workflow with an authorized face, permitted speech, and a careful performance review.

Combine a clear portrait with an authorized audio performance. The connected workflow uses the portrait as its starting-image input and the supplied recording for the performance. Inspect the resulting frames and audio rather than assuming perfect identity or lip synchronization.

Explore Video Generator

Talking Face / Photo + audioRecorded Studio result
Actual five-second Studio result from a synthetic portrait and an original generated recording. Both source files are included in the recipe. Compare the mouth, eyes and timing with the source audio; the requested speech script has not been independently transcribed.View the recorded recipe
01 / Start with
Authorized photo + audio
02 / Direct
Appearance and performance
03 / Review
A talking-face video

From input to intention

Make each input’s role unmistakable.

Prepare the input, express the intention, and review what actually changed.

01

Prepare a clear, permitted photo

Choose an image in which the face and important features are readable. Confirm that the person has authorized the intended animated use and context.

02

Choose the audio performance

Use permitted speech with the intended wording and delivery. Keep a transcript so timing and meaning can be checked against the eventual result.

03

Review the complete performance

Watch the mouth, eyes, facial boundaries, head movement, and relationship to the recording. Listen for the exact wording and reject misleading or unapproved uses.

Start with a useful brief

The recorded performance direction.

This is the actual submitted prompt for the sample above.

Example direction

Frame alignment: Picture 1 is the exact opening portrait at zero seconds. integrated_multimodal_description: [Shot 1] The original fictional adult woman in the supplied portrait faces the camera in the same softly lit room. Preserve her face, hair, clothes and background throughout a five-second chest-up shot. Her lips and jaw follow the supplied original speech recording naturally, with restrained head movement, subtle blinking and a calm friendly expression. Keep the camera locked and finish with a natural settled expression. No cuts, new people, text or watermark. overall_soundscape: Reuse the supplied audio performance unchanged. Do not add voices or sound effects. non_diegetic_music: N/A

Exact recorded prompt for this five-second Studio output. Open the recipe for both source files and all settings; listen to compare the performance with the original recording.

01

Photo plus audio is the clear input label

The connected Talking Face workflow supplies the portrait as the starting-image input and reuses the supplied recording. This differs from generating a new scene inspired by appearance or voice references.

The starting image is an input constraint, not a promise that every later facial detail remains unchanged. Inspect the complete animation and compare it with the source.

02

Check permission for both sides of the performance

The person in the photo and the speaker in the audio may not be the same. Make sure the intended combination is authorized and does not falsely attribute statements or actions.

Do not use this workflow for deceptive impersonation, fraudulent endorsements, harassment, or non-consensual intimate material. Keep the approved script and source permissions with the project.

03

Inspect more than the mouth

Lip movement matters, but the rest of the face also contributes to the performance. Watch the eyes, jaw, hairline, and any head movement for unwanted distortion or changes in identity.

Review pauses, fast syllables, and the beginning and ending. A convincing short moment does not establish that the entire performance is usable.

04

A real demonstration needs the full chain

This example includes the actual synthetic portrait, original generated audio, exact direction, detailed rendering controls and generated clip. No real person’s voice was cloned.

The speech prompt contains the requested words, not an independently verified transcript. Listen to the recording and review the complete video before using it. Opening its recipe leaves consent unchecked for your intended use.

The useful details

A few good
questions.

Where does Talking Face live?

Video Generator → Video Generator → Talking Face. It is a use case inside a broad mode, not a new main-sidebar tool.

Is the photo guaranteed to be the first frame?

The connected workflow uses your portrait as its starting-image input. Review the actual output: conditioning does not guarantee pixel-identical frames or unchanged identity throughout the performance.

Does the photo and audio have to be the same person?

Use a combination that is permitted and does not misrepresent an identity or statement. The final supported combinations depend on the connected workflow.

Can I generate a talking-face video here?

Open this workflow in Studio to prepare its inputs, direction, and project assets. Review the supported options and the service quote before generating.

Can I animate singing instead of speech?

Singing Characters is a related photo-and-audio use case under Video Generator. It focuses on sustained notes, breath, expression, and synchronization with a permitted vocal performance.

Connected workflows

A different starting point?

All use cases

One tool. More possibilities.

One reference mode. A clear use case.

This use case belongs inside the broader tool. Explore related approaches without learning a different navigation system.

Open in Studio