Create a [duration] [format] in [language]. Scene: [place, time, mood, audience]. Voice: [speaker count, traits, emotion, pace]. Dialogue: "[exact words or conversation outline]". Ambience: [room tone, weather, crowd, environment]. Music: [style, instruments, intensity, entry or fade]. Sound effects in order: [event 1], [event 2], [ending cue]. Keep the result [clean / cinematic / broadcast-ready].

Write prompts like compact sound production briefs.
Turn a creative idea into instructions Seed Audio 2.0 can follow. Structure the scene, map references, direct each voice, and place music, ambience, and effects in playback order.
A strong prompt is a timeline, not a tag list.
Seed Audio 2.0 can combine speech, character delivery, ambience, music, and foley-style events. Guide the mix as a short production brief in playback order: first the scene, then the speakers, then the sound layers, then the ending.
Start from the right template, then remove anything unnecessary.
Use @audio1 for [speaker or voice trait]. Use @audio2 for [second speaker or style reference]. Create a [duration] [format] in [language]. Speaker map: [speaker A follows @audio1], [speaker B follows @audio2]. Scene: [place, mood, purpose]. Delivery: [emotion, rhythm, accent, pauses]. Dialogue: "[lines or conversation structure]". Layers: [ambience], [music bed], [sound effects in order]. Keep voices consistent and speech easy to understand.
Build the brief from four decisions.
Scene
Name the format, setting, audience, language, mood, and rough length before writing dialogue.
Voice
Define speaker count, vocal texture, accent, emotion, pacing, and the clarity priority.
Timeline
Write the Seed Audio 2.0 prompt in playback order so speech, ambience, music, and events do not compete.
Layers
Add one ambience bed, one music direction, and only the sound effects that matter to the scene.
Choose the prompt mode before writing the first line.
The content may look similar, but the job is different. T2A asks Seed Audio 2.0 to invent a voice and scene. TA2A asks it to follow reference signals while still obeying your written direction.
T2A: text or image to audio
Use this when you want Seed Audio 2.0 to invent the voice and acoustic scene from a written brief, optionally guided by one image reference.
TA2A: text plus audio references
Use this when voice identity, pacing, emotional texture, or multi-speaker consistency matters. Map every reference clearly inside the prompt.

Describe what happens in the order the listener hears it.
“Create a 25-second cinematic podcast intro in English. Begin in a quiet midnight radio studio with soft rain against the window. A warm, calm female host says, ‘Every city has a frequency after dark.’ Add a low analog synth pulse after the final word, then distant traffic, one tape-button click, and a clean two-second fade.”
Put music, ambience, and effects where they happen.
Time words help Seed Audio 2.0 understand priority. Use them to stage the mix, not to micromanage every frame.
Opening cue
Establish room tone, weather, a phone buzz, door movement, or a simple musical entrance.
Main speech
Place dialogue or narration here with the most important voice and emotion directions.
Scene movement
Add one or two story-supporting sound events without covering the voice.
Exit
End with a fade, resolved chord, bell, door, breath, button sound, or another clean cue.
Three reusable Seed Audio 2.0 prompt shapes.
Use these as structures, then replace the scene, speakers, references, and sound events with your own creative direction.

Cinematic phone call
Create a 45-second cinematic crime scene in English. Start with a quiet phone vibration, distant rain, and a low suspense pad. Detective A speaks in a restrained, tired voice, close to the microphone. Detective B answers through a slightly distorted phone line with a calm but evasive tone. Keep the dialogue clear. Add two small footsteps and one car brake near the end.
Reference-guided livestream
Use @audio1 for Host A and @audio2 for Host B. Create a 60-second energetic livestream shopping segment in English. Host A is bright, fast, and persuasive. Host B is warmer and reacts with short supportive lines. Add very light upbeat background music, subtle package handling sounds, and a clean call-to-action ending.
Image-based fantasy narrator
Use @image as the scene reference. Create a 35-second fantasy narration in English. The speaker is a serious middle-aged astrologer with slow, ceremonial delivery and a deep resonant voice. Add quiet rain outside a stone observatory, one soft page turn, and a distant bell. Keep the music minimal and mysterious.
If the output feels messy, simplify the brief first.
Listing too many layers
Five music moods, ten effects, and several character arcs usually create a muddy mix. Reduce the layer count first.
Forgetting the reference map
For TA2A, say exactly which character or style should follow @audio1, @audio2, or @audio3.
Writing a keyword cloud
Seed Audio 2.0 works better when the prompt reads like a compact production brief in playback order.
Overacting the emotion
Extreme emotional adjectives can make delivery less natural. Begin with one controlled cue.
Treating duration as exact
Requested length is guidance. Leave room for trimming when the final edit needs frame-accurate timing.
Using protected names
Describe roles and vocal qualities generically instead of naming real people or protected characters.
Use the guide with real Seed Audio 2.0 examples.
Start with these templates, then test them in the live workspace and compare how small prompt changes reshape the generated scene.