Video / Field guide

A clear photo. An approved performance.

Prepare an authorized portrait, permitted audio, and a transcript for the Talking Face workflow.

DeepAny workflow guide 3 min readPractical workflow
Talking Face / Photo + audioRecorded Studio result
Actual five-second Studio result from a synthetic portrait and an original generated recording. Both source files are included in the recipe. Compare the mouth, eyes and timing with the source audio; the requested speech script has not been independently transcribed.View the recorded recipe

Talking Face combines an appearance reference with an audio performance. Before judging animation quality, make the roles and permissions of those inputs clear.

01 / Video

Define the intended identity and statement

Confirm who appears in the photo and whose performance appears in the recording. Check that the intended combination and publication context are authorized.

Do not treat a publicly available portrait or voice clip as a license for a new performance. Keep the approved purpose and wording with the project.

02 / Video

Choose a readable face reference

Use a photo in which the relevant facial features can be inspected. Note any hair, glasses, hands, or strong shadows that may complicate the intended animation.

Follow the connected mode’s actual input requirements for the request. This guide does not prescribe unverified pixel dimensions or an exact supported pose range.

03 / Video

Prepare the audio and transcript

Choose the approved recording and keep an accurate transcript. Listen for background noise, abrupt starts or endings, and unclear wording before using it as a performance reference.

Separate the script from the direction. A note such as “calm delivery” should not accidentally become part of the spoken content.

04 / Video

Use the photo label accurately

The connected Talking Face workflow supplies the portrait as the starting-image input and reuses the supplied recording. This differs from generating a new scene inspired by appearance or voice references.

The starting image is an input constraint, not a promise that every later facial detail remains unchanged. Inspect the complete animation and compare it with the source.

05 / Video

Review the full performance

Watch the mouth, eyes, head movement, and facial boundaries while listening to the reference. Review pauses, rapid speech, and the beginning and ending.

Compare the wording with the transcript and the visual identity with the photo. A convincing fragment does not establish that the whole performance is accurate or appropriate.

06 / Video

Keep the demonstration honest

This example includes the actual synthetic portrait, original generated audio, exact direction, detailed rendering controls and generated clip. No real person’s voice was cloned.

The speech prompt contains the requested words, not an independently verified transcript. Listen to the recording and review the complete video before using it. Opening its recipe leaves consent unchecked for your intended use.

A direction you can adapt

Example direction

Frame alignment: Picture 1 is the exact opening portrait at zero seconds. integrated_multimodal_description: [Shot 1] The original fictional adult woman in the supplied portrait faces the camera in the same softly lit room. Preserve her face, hair, clothes and background throughout a five-second chest-up shot. Her lips and jaw follow the supplied original speech recording naturally, with restrained head movement, subtle blinking and a calm friendly expression. Keep the camera locked and finish with a natural settled expression. No cuts, new people, text or watermark. overall_soundscape: Reuse the supplied audio performance unchanged. Do not add voices or sound effects. non_diegetic_music: N/A

Exact recorded prompt for this five-second Studio output. Open the recipe for both source files and all settings; listen to compare the performance with the original recording.

Keep this close

A brief worth keeping.

  • Authorize the identity and the statement together.
  • Use a readable portrait and an approved recording.
  • Keep a transcript for review.
  • Evaluate the whole performance, not one good moment.

Connected workflows

Continue your practice.

All use cases