Define the listening context
Decide whether the audio needs to stand alone or fit a video, presentation, or other existing material. Keep the intended role clear in the brief.
Audio / Audio Generator
Plan an audio direction from text, with a brief that separates the sound, delivery, and intended use.
Begin with what the listener should hear: a spoken message, a particular delivery, or a sound that supports a scene. Make the purpose, pacing, and ending clear before evaluating the result by listening, not by looking at a waveform.
Your inputs, directions, and versions stay together in Studio.
Listen to the morning breeze.
Choose your starting point
Generation, voice cloning, and continuation have different inputs and should remain easy to distinguish.
A considered workflow
A useful evaluation happens through listening, not by looking at the player.
Decide whether the audio needs to stand alone or fit a video, presentation, or other existing material. Keep the intended role clear in the brief.
Specify the desired delivery, pace, texture, and development where the connected mode supports those instructions. Do not assume every type of audio uses the same controls.
Check clarity, distracting artifacts, and whether the result fits its destination. For speech, verify pronunciation and wording rather than trusting a waveform.
Make the brief do the work
This is the actual brief for the playable result, not a separate speech example.
Soft birdsong and leaves rustling in a gentle morning breeze. No speech, no music.
Exact prompt submitted for this recording. The matching library recipe restores the five-second duration, Euler sampler, workflow-default schedule and fixed seed.
Speech, music, ambience, and effects create different expectations. The final connected modes should say clearly which kinds of input and output they support.
Choose the audio type and write your direction in Studio. The connected service must confirm sound types, languages, duration limits, and licensing; a written example is not evidence of a generated result.
A useful demonstration includes the exact brief, the actual recording, and an explanation of what to listen for. For speech, add a transcript. For comparisons, keep the relevant source available.
This page includes the actual Studio recording and its submitted prompt. The recorded recipe links the file to its duration, sampling controls and seed.
Creating new audio from a brief and imitating a particular voice are not the same task. A voice-specific workflow must begin with authorization to use the source.
Use the separate Voice Cloning overview for that intention. The public navigation remains broad while the use-case pages explain the distinct preparation and review process.
The useful details
Open this workflow in Studio to prepare its inputs, direction, and project assets. Review the supported options and the service quote before generating.
You can play the published recording and load the recorded setup in Studio. The prompt, duration, sampling controls and seed are included; backend changes or nondeterministic processing can still change a fresh result.
No. Voice cloning uses a particular consented reference voice. Text-led generation does not by itself establish a right to use someone’s vocal identity.
A five-second, 32 kHz stereo FLAC file. FLAC keeps the generated audio lossless. The recorded setup uses the connected Text to Audio workflow.
Connected workflows
Your next step
Explore the possibilities here, then continue in the existing Studio.