结论先说:官方两份指南(
VIDEO_PROMPT_WRITING_GUIDE_base_en.md/..._ref_en.md) 没有「lip sync / 对口型」这一节,也没有任何"让嘴唇对上"的专门句式。 官方给出的唯一手段是:把逐句台词/歌词原样写进<d>[语言] …</d>、给发声者稳定 ID(Sx)、 并用<Audio N>+ 复制关系标注音频来源。你现在用的 Director 模板里一个<d>都没有 —— 模型拿不到"说了什么/唱了什么"的时间线,嘴型只能自由发挥,这就是没对上的直接原因。
来源(两者内容一致,已验证与官方 diff 仅差一个空行):
skills/h3-prompt-writing/references/)<d>(base §4.4 Speakers, Dialogue, and Singing)Subjects who speak, sing, or produce an off-screen human voice use stable IDs such as
(S1)and(S2). When multiple already-numbered speakers speak or sing together, use a compound ID such as(S1,S2). A speaker keeps the same ID across shots; characters who never vocalize receive no speaker ID.When a speaker first appears, provide enough information from the visual and audio context to establish a stable identity, such as character type, age, gender, whether the person is on-screen, pitch, timbre, speaking rate, or accent. Place the speaker's identifying phrase, ID, action, and delivery outside
<d>. Inside<d>, include only the language tag and the actual user-provided spoken content. Preserve every original word and punctuation mark verbatim; do not translate or rewrite them.
示例(官方):
The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>
For voiceover, use the exact phrase
says in an off-screen voiceover. Immediately after every voiceover<d>block, state that the corresponding on-screen character's lips remain closed:
The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.
When the same line of dialogue or lyrics crosses a cut, use
<scenetrans>at the connecting points in both parts and explicitly state that the audio continues across the cut. Use<cutoff>when speech is truncated by the end of the video. Continuity may be expressed withcontinues seamlessly across the cut,continues uninterrupted into the next shot,carries over from the previous shot, orremains audible across the transition.
When a referenced subject physically speaks, retain both the visual reference label and the speaker ID:
<Subject 2> (S1) turns toward the woman and says, <d>[English] Last summer, I went to my grandfather's house. He talked about you.</d>
<Subject N>identifies the referenced subject, while(Sx)identifies the actual speaker. When the subject speaks, write<Subject N> (Sx). If the same subject speaks off-screen, keep the same form and mark it asoff-screen.Assign
(Sx)once according to the order of actual vocal events in the target video. Reuse the corresponding ID at every actual vocal event indetailed_description… Do not write(Sx)inretention_analysis.When dialogue, narration, or lyrics from reference audio are directly reused, or when the input prompt explicitly requests their reperformance, preserve the exact source words and original language inside
<d>. Write[unclear]for unintelligible spans instead of guessing or paraphrasing them. Standardize punctuation to the basic written marks needed to express the sentence, such as,,.,?, and!.When only timbre, rhythm, emotion, or delivery is referenced, do not carry the original dialogue from the reference audio into the target video.
| Relationship marker | Meaning(官方表) |
|---|---|
fully_copy | The complete source audio serves as the target video's complete final audio track |
partially_copy | Only part of the timeline or selected audio layers are copied, or other sounds are added, removed, or replaced after copying |
reference | The signal is not copied directly; only timbre, rhythm, music style, dialogue content, or sound texture is referenced |
weak_reference | Only broad similarity in category or atmosphere is retained |
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
<Audio N>represents a standalone audio asset or an enabled synchronized audio track from a reference video.<Audio 2> is the synchronized audio track of <Video 1> and is reused in the target video.
ref §5.2 表格:Audio relationships — "Cites <Audio N> in the corresponding shot or audio phase and states whether the signal is copied or referenced"。
ref §6:Write complete dialogue and lyrics only inside <d> in detailed_description; do not repeat them in these two sections(指 overall_soundscape / non_diegetic_music);环境音复制关系写在 overall_soundscape,观众专属配乐关系写在 non_diegetic_music。
官方 Ref2VA 示例原文里紧接台词之后写的是嘴部状态:
She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
即:台词用 <d>,嘴部后续动作/开合状态用普通英文散文描述,官方没有任何"lip sync""mouth matches audio"之类的关键词写法。
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.
<Audio 1>: partially_copy - the source vocal layer and ambience are reused, and the target video's own ambience is layered on top.
description 里(每个发声点都写):
[Shot 1] … <Subject 2> (S1) sings, <d>[Chinese] 第一句歌词按原样逐字写在这里。</d> …
[Shot 2] At 00:08.000, … the seven characters (S1,S2,S3,S4,S5,S6,S7) sing together, <d>[Chinese] 副歌歌词按原样逐字写在这里。</d> …
<Subject 2> (S1) says, <d>[Chinese] 台词原样。</d>
<Subject 2> (S1) says in an off-screen voiceover: <d>[Chinese] 旁白原样。</d> while their lips remain completely closed.
[Shot 2] At 00:06.000, <scenetrans> the same line continues seamlessly across the cut … <d>[Chinese] 后半句。</d>
[Shot 3] At 00:12.000, <cutoff> the speech is truncated by the end of the video.
你现在的模板(Replace the person in <Video 1> …)保留,只在其后追加两段即可:
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.
<Audio 1>: partially_copy - the source vocal layer is reused; the lip movement of each singing character must follow the lyrics written in <d> blocks.
注意:官方没有"lip movement must follow …"这种句子——这句是我为表达意图加的,不是官方句式; 官方路线是"把 <d> 台词写全 + <Audio N> 标注关系",靠这两件事让口型有依据。要不要用这句由你定。
lip sync / sync lips / mouth matches audio / viseme 这类关键词条款,官方指南里搜不到。<d>(官方要求 verbatim,只保留 , . ? !,去掉重复波浪号/emoji/装饰标点)。[unclear],不要猜。<d> 不要写进 retention_analysis,也不要写进 overall_soundscape / non_diegetic_music。(Sx)。