两个事实叠加:
ComfyUI_MiniMaxH3_Director/nodes/conditioning.py 在 rv2v 时走 MiniMaxH3ReferenceToVideo,把参考视频(含其音轨,经 audio_vae)作为条件信号。 所以模型"听得到"这首歌的声学信息(节奏、音色、时序)。director/audio_export.py AUDIO_MODE_SOURCE,且 VIDEO_EDIT_AUDIO_TASKS = {v2v, rv2v, r2v} 包含 rv2v → 输出音频直接采用源视频音轨(passthrough),不是模型自己生成的那条。结论:嘴是模型按它自己"想象"的唱词生成的,音轨却是原唱的。你现在的提示词里没有一句 <d> 台词/歌词,模型没有任何文字层的唱词锚点,只能凭声学信号猜 → 嘴型漂。 官方给的唯一手段就是:把原唱词逐字写进 <d>,让模型的"想象唱词"等于原唱唱词。
subject_definitions 里加音频定义<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.
retention_analysis 里加(二选一,看你的 audioMode 实际行为):
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
<Audio 1>: partially_copy - only the source vocal layer is copied; the target video's other sounds are generated.
官方对四个标记的定义(ref §4.2):
| marker | 含义 |
|---|---|
fully_copy | 完整源音频作为目标视频的完整最终音轨 |
partially_copy | 只复制部分时间线/部分层,或复制后又增删替换了别的声 |
reference | 不直接复制,只参考音色、节奏、音乐风格、台词内容或声音质感 |
weak_reference | 只保留类别/氛围上的大致相似 |
⚠️ 官方 ref §5.4 有一条关键限制:"When only timbre, rhythm, emotion, or delivery is referenced, do not carry the original dialogue from the reference audio into the target video." → 如果你把音频标成 reference,就不该把原歌词带进来;想对上原唱,必须走 copy 关系 + 把歌词写进 <d>。
detailed_description 里逐句写唱词单人主唱:
[Shot 1] … <Subject 2> (S1) sings, <d>[Chinese] 第一句原样逐字。</d> …
[Shot 2] At 00:06.000, … <Subject 2> (S1) sings, <d>[Chinese] 第二句原样逐字。</d> …
七人齐唱(手势舞最常见):用复合 ID,官方 base §4.4 明确允许 ("When multiple already-numbered speakers speak or sing together, use a compound ID such as (S1,S2)"):
[Shot 1] … the seven characters (S1,S2,S3,S4,S5,S6,S7) sing together, <d>[Chinese] 齐唱那句原样逐字。</d> …
跨段/跨剪切接续(36.2s 必须切 ≥3 段,每段 ≤362 帧):
[Shot 3] At 00:09.000, <scenetrans> the same vocal line continues seamlessly across the cut, <d>[Chinese] 后半句原样逐字。</d>
被结尾截断:
<cutoff> the vocal line is truncated by the end of the video.
<d> 内只放"语言标签 + 逐字原唱词",不翻译、不改写、不意译; 标点只保留 , . ? !,去掉重复波浪号、emoji、装饰标点。[unclear],不要猜、不要改写(官方 ref §5.4 原文:instead of guessing or paraphrasing them)。<d> 外面,<d> 里只有语言标签和内容。(Sx) 按实际发声顺序分配一次,之后每个发声点复用;不发声的角色不给 ID; 不要把 (Sx) 写进 retention_analysis。<d> 只写在 detailed_description;不要写进 retention_analysis,也不要在 overall_soundscape / non_diegetic_music 里重复歌词(官方 ref §6)。lip sync / sync lips / mouth matches audio / viseme 之类官方关键词,官方两份指南里搜不到—— 写了也没用,别自造。官方要求 verbatim,所以这段必须落在真实歌曲上,我不能编。两条路:
(Sx) 和时间戳,直接产出可粘贴版本;素材:GPU 机 ComfyUI/input/share_93e5d5fc2dc0eac728abda226f95c5431788950922652-无片尾.mp4 (1280×720 / 30fps / 1086 视频帧 / 36.2s)。
skills/h3-prompt-writing/references/,与本地副本已 diff 核对一致)ComfyUI/custom_nodes/ComfyUI_MiniMaxH3_Director/nodes/conditioning.py、director/audio_export.py