# 手势舞 + 唱歌：口型修正补丁（2026-09-14）

## 一、先说为什么你的嘴型会对不上（你这台机器的实际代码 + 官方文档）

两个事实叠加：

1. **源音频确实喂进了模型**：`ComfyUI_MiniMaxH3_Director/nodes/conditioning.py`
   在 rv2v 时走 `MiniMaxH3ReferenceToVideo`，把参考视频（含其音轨，经 `audio_vae`）作为条件信号。
   所以模型"听得到"这首歌的声学信息（节奏、音色、时序）。
2. **但最终音轨是原唱**：`director/audio_export.py`
   `AUDIO_MODE_SOURCE`，且 `VIDEO_EDIT_AUDIO_TASKS = {v2v, rv2v, r2v}` 包含 rv2v
   → 输出音频直接采用源视频音轨（passthrough），**不是**模型自己生成的那条。

结论：**嘴是模型按它自己"想象"的唱词生成的，音轨却是原唱的**。你现在的提示词里没有一句
`<d>` 台词/歌词，模型没有任何文字层的唱词锚点，只能凭声学信号猜 → 嘴型漂。
官方给的唯一手段就是：**把原唱词逐字写进 `<d>`，让模型的"想象唱词"等于原唱唱词**。

## 二、补丁（只加这两块，动作迁移部分一个字都不用改）

### 1）在 `subject_definitions` 里加音频定义

```text
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.
```

`retention_analysis` 里加（二选一，看你的 audioMode 实际行为）：

```text
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
```
```text
<Audio 1>: partially_copy - only the source vocal layer is copied; the target video's other sounds are generated.
```

官方对四个标记的定义（ref §4.2）：

| marker | 含义 |
|---|---|
| `fully_copy` | 完整源音频作为目标视频的完整最终音轨 |
| `partially_copy` | 只复制部分时间线/部分层，或复制后又增删替换了别的声 |
| `reference` | 不直接复制，只参考音色、节奏、音乐风格、台词内容或声音质感 |
| `weak_reference` | 只保留类别/氛围上的大致相似 |

⚠️ 官方 ref §5.4 有一条关键限制：**"When only timbre, rhythm, emotion, or delivery is referenced,
do not carry the original dialogue from the reference audio into the target video."**
→ 如果你把音频标成 `reference`，就**不该**把原歌词带进来；**想对上原唱，必须走 copy 关系 + 把歌词写进 `<d>`。**

### 2）在 `detailed_description` 里逐句写唱词

单人主唱：

```text
[Shot 1] … <Subject 2> (S1) sings, <d>[Chinese] 第一句原样逐字。</d> …
[Shot 2] At 00:06.000, … <Subject 2> (S1) sings, <d>[Chinese] 第二句原样逐字。</d> …
```

七人齐唱（手势舞最常见）：用**复合 ID**，官方 base §4.4 明确允许
（"When multiple already-numbered speakers speak or sing together, use a compound ID such as `(S1,S2)`"）：

```text
[Shot 1] … the seven characters (S1,S2,S3,S4,S5,S6,S7) sing together, <d>[Chinese] 齐唱那句原样逐字。</d> …
```

跨段/跨剪切接续（36.2s 必须切 ≥3 段，每段 ≤362 帧）：

```text
[Shot 3] At 00:09.000, <scenetrans> the same vocal line continues seamlessly across the cut, <d>[Chinese] 后半句原样逐字。</d>
```
被结尾截断：
```text
<cutoff> the vocal line is truncated by the end of the video.
```

## 三、官方铁律（违反就会被模型忽略或写坏）

1. **verbatim**：`<d>` 内只放"语言标签 + 逐字原唱词"，**不翻译、不改写、不意译**；
   标点只保留 `,` `.` `?` `!`，去掉重复波浪号、emoji、装饰标点。
2. **听不清写 `[unclear]`**，不要猜、不要改写（官方 ref §5.4 原文：`instead of guessing or paraphrasing them`）。
3. 说话人的**身份描述、ID、动作、唱法放在 `<d>` 外面**，`<d>` 里只有语言标签和内容。
4. `(Sx)` **按实际发声顺序分配一次**，之后每个发声点复用；不发声的角色不给 ID；
   **不要**把 `(Sx)` 写进 `retention_analysis`。
5. `<d>` 只写在 `detailed_description`；**不要**写进 `retention_analysis`，也**不要**在
   `overall_soundscape` / `non_diegetic_music` 里重复歌词（官方 ref §6）。
6. 台词密集时官方要求**优先保证完整唱词时间线**，别为凑字数删句（ref §5.2）。
7. **不存在** `lip sync` / `sync lips` / `mouth matches audio` / `viseme` 之类官方关键词，官方两份指南里搜不到——
   写了也没用，别自造。

## 四、还差一样东西：真实歌词

官方要求 verbatim，所以这段**必须落在真实歌曲上**，我不能编。两条路：
1. 你把歌词贴给我（或告诉我歌名），我按段拆句、配 `(Sx)` 和时间戳，直接产出可粘贴版本；
2. 我用 ASR 从源视频音轨听写——**两台机器目前都没装 ASR**（GPU 机与 DSH 机都没有 whisper / faster-whisper / funasr），
   要装（pip + 模型下载）才能跑，且听写结果仍建议你核对一遍。

素材：GPU 机 `ComfyUI/input/share_93e5d5fc2dc0eac728abda226f95c5431788950922652-无片尾.mp4`
（1280×720 / 30fps / 1086 视频帧 / 36.2s）。

## 五、出处

- 官方 base 指南 https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
- 官方 ref 指南 https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
- 官方仓库 https://github.com/MiniMax-AI/MiniMax-H3 （`skills/h3-prompt-writing/references/`，与本地副本已 diff 核对一致）
- 本机节点代码：`ComfyUI/custom_nodes/ComfyUI_MiniMaxH3_Director/nodes/conditioning.py`、`director/audio_export.py`
