H3口型台词写法-官方条款与补丁-20260914.md

H3 口型/台词写法:官方条款 + 补丁(2026-09-14)

结论先说:官方两份指南(VIDEO_PROMPT_WRITING_GUIDE_base_en.md / ..._ref_en.md) 没有「lip sync / 对口型」这一节,也没有任何"让嘴唇对上"的专门句式。 官方给出的唯一手段是:把逐句台词/歌词原样写进 <d>[语言] …</d>、给发声者稳定 ID (Sx)、 并用 <Audio N> + 复制关系标注音频来源。你现在用的 Director 模板里一个 <d> 都没有 —— 模型拿不到"说了什么/唱了什么"的时间线,嘴型只能自由发挥,这就是没对上的直接原因。

一、官方原文条款(照抄,未改写)

来源(两者内容一致,已验证与官方 diff 仅差一个空行):

1)说话人 ID 与 <d>(base §4.4 Speakers, Dialogue, and Singing)

Subjects who speak, sing, or produce an off-screen human voice use stable IDs such as (S1) and (S2). When multiple already-numbered speakers speak or sing together, use a compound ID such as (S1,S2). A speaker keeps the same ID across shots; characters who never vocalize receive no speaker ID.

When a speaker first appears, provide enough information from the visual and audio context to establish a stable identity, such as character type, age, gender, whether the person is on-screen, pitch, timbre, speaking rate, or accent. Place the speaker's identifying phrase, ID, action, and delivery outside <d>. Inside <d>, include only the language tag and the actual user-provided spoken content. Preserve every original word and punctuation mark verbatim; do not translate or rewrite them.

示例(官方):

The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>

2)唯一直接提嘴唇的硬规则:旁白(base §4.4)

For voiceover, use the exact phrase says in an off-screen voiceover. Immediately after every voiceover <d> block, state that the corresponding on-screen character's lips remain closed:

The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.

3)跨剪切/被结尾截断(base §4.4)

When the same line of dialogue or lyrics crosses a cut, use <scenetrans> at the connecting points in both parts and explicitly state that the audio continues across the cut. Use <cutoff> when speech is truncated by the end of the video. Continuity may be expressed with continues seamlessly across the cut, continues uninterrupted into the next shot, carries over from the previous shot, or remains audible across the transition.

4)Ref2VA 里被引用主体开口的形式(ref §5.4)

When a referenced subject physically speaks, retain both the visual reference label and the speaker ID:

<Subject 2> (S1) turns toward the woman and says, <d>[English] Last summer, I went to my grandfather's house. He talked about you.</d>

<Subject N> identifies the referenced subject, while (Sx) identifies the actual speaker. When the subject speaks, write <Subject N> (Sx). If the same subject speaks off-screen, keep the same form and mark it as off-screen.

Assign (Sx) once according to the order of actual vocal events in the target video. Reuse the corresponding ID at every actual vocal event in detailed_description … Do not write (Sx) in retention_analysis.

When dialogue, narration, or lyrics from reference audio are directly reused, or when the input prompt explicitly requests their reperformance, preserve the exact source words and original language inside <d>. Write [unclear] for unintelligible spans instead of guessing or paraphrasing them. Standardize punctuation to the basic written marks needed to express the sentence, such as ,, ., ?, and !.

When only timbre, rhythm, emotion, or delivery is referenced, do not carry the original dialogue from the reference audio into the target video.

5)音频来源标注(ref §4.2 Audio + §5.2)

Relationship markerMeaning(官方表)
fully_copyThe complete source audio serves as the target video's complete final audio track
partially_copyOnly part of the timeline or selected audio layers are copied, or other sounds are added, removed, or replaced after copying
referenceThe signal is not copied directly; only timbre, rhythm, music style, dialogue content, or sound texture is referenced
weak_referenceOnly broad similarity in category or atmosphere is retained
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.

<Audio N> represents a standalone audio asset or an enabled synchronized audio track from a reference video. <Audio 2> is the synchronized audio track of <Video 1> and is reused in the target video.

ref §5.2 表格:Audio relationships — "Cites <Audio N> in the corresponding shot or audio phase and states whether the signal is copied or referenced"。

ref §6:Write complete dialogue and lyrics only inside <d> in detailed_description; do not repeat them in these two sections(指 overall_soundscape / non_diegetic_music);环境音复制关系写在 overall_soundscape,观众专属配乐关系写在 non_diegetic_music。

6)官方示例里"嘴部动作"是散文写的(不是关键词)

官方 Ref2VA 示例原文里紧接台词之后写的是嘴部状态:

She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
He closes his mouth into an apologetic smile and strokes the dog's thick white fur.

即:台词用 <d>,嘴部后续动作/开合状态用普通英文散文描述,官方没有任何"lip sync""mouth matches audio"之类的关键词写法。

7)音频输入上限(官方 README,Ref2VA 输入表)

二、可以直接粘贴的补丁

A. 唱歌 / 群唱(你这种多人唱跳最可能是这种)

<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.
<Audio 1>: partially_copy - the source vocal layer and ambience are reused, and the target video's own ambience is layered on top.

description 里(每个发声点都写):

[Shot 1] … <Subject 2> (S1) sings, <d>[Chinese] 第一句歌词按原样逐字写在这里。</d> …
[Shot 2] At 00:08.000, … the seven characters (S1,S2,S3,S4,S5,S6,S7) sing together, <d>[Chinese] 副歌歌词按原样逐字写在这里。</d> …

B. 单人说话(被引用主体开口)

<Subject 2> (S1) says, <d>[Chinese] 台词原样。</d>

C. 旁白(画外音)

<Subject 2> (S1) says in an off-screen voiceover: <d>[Chinese] 旁白原样。</d> while their lips remain completely closed.

D. 跨剪切 / 被结尾截断

[Shot 2] At 00:06.000, <scenetrans> the same line continues seamlessly across the cut … <d>[Chinese] 后半句。</d>
[Shot 3] At 00:12.000, <cutoff> the speech is truncated by the end of the video.

E. 加到你现在那个模板上的最小改法

你现在的模板(Replace the person in <Video 1> …)保留,只在其后追加两段即可:

<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.
<Audio 1>: partially_copy - the source vocal layer is reused; the lip movement of each singing character must follow the lyrics written in <d> blocks.

注意:官方没有"lip movement must follow …"这种句子——这句是我为表达意图加的,不是官方句式; 官方路线是"把 <d> 台词写全 + <Audio N> 标注关系",靠这两件事让口型有依据。要不要用这句由你定。

三、不要写的东西(官方文档里不存在)

下载此文件