报告_MiniMax_H3_REF2VA_动作归属问题.md

MiniMax H3 / Ref2VA 动作归属问题调研报告

场景: ComfyUI Dasiwa 工作流,REF2VA(参考图生视频)。参考图是浅蓝浴巾亚洲女生三视图。提示词要求第一人称站在门外,镜头侧一只手握住门把手下压推开磨砂玻璃门;参考图女生(<Subject 1>)全程站在房间深处花洒下方,双手抓住胸前浴巾顶部,不靠近门、不碰门、不做开门动作。

实测失败: 9 个版本迭代后,Subject 1 仍「同步执行开门动作」——门开时她的手出现在门框/门把手上,或她走到门边扶门。开门被模型当成她在做,而不是 POV 镜头侧那只手。

已试过且仍失败: 正面表述(无否定词)、官方六段式(subject_definitions / summary / retention_analysis / detailed_description / overall_soundscape / non_diegetic_music)。

调研范围: MiniMax 官方 prompt guide、ComfyUI 原生文档与 GitHub issues、ethanfel/ComfyUI-MiniMax-H3-Guide、teskor-hub/minimax-h3-skill、Reddit r/StableDiffusion、Kapwing / APIDot / Hailuo 中文指南、Fun ControlNet / Add Guide 工作流资料。截止日期约 2026-09-02。

结论先行: 这不是六段式写错那么简单。它是 Ref2VA 的结构性偏向:参考图人物是画面里唯一有视觉 token 的「施事者」,提示词里任何显著肢体动作(握手、下压、推门)都会默认落到这个人身上。官方 fully_preserved 不锁姿态/站位,官方运镜词 POV 的字面意思是「主体视角」而不是「门外第三人称手」。社区没有一篇 issue 用完全相同的标题描述「参考角色被拉进镜头/第二人的动作」,但同一机制在多处被反复踩到,解法也已经收敛。


1. 这是不是已知痛点?官方和社区怎么解释?

1.1 没有完全同名的「官方 bug 标题」

在 Comfy-Org/ComfyUI issues、MiniMax 官方 guide、ethanfel 仓库里,没有找到一条标题等价于:

reference character gets pulled into actions described for the camera / POV / second person

所以不宜把它说成「官方已确认的 named bug」。更准确的说法是:这是 Ref2VA 默认施事者绑定 + 官方 retention 语义 + 官方 POV 词义,三者叠加后的可复现失败模式。 同类现象在社区已经被写成可操作的 troubleshooting。

1.2 官方机制:fully_preserved 保的是「已定义角色」,不是「站着别动」

MiniMax 全参考模式指南把 <Subject N> 定义为「会实际出现在目标视频里的可复用可见内容」,范围包括人、物、场景、服装,也包括动作、表情、姿态:

<Subject N> is used for reusable visible content, including: People, animals, or objects; Scenes, backgrounds, or environments; Clothing, props, interfaces, or visual effects; Styles, actions, expressions, or poses.

出处:https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md 镜像:https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/h3-prompt-writing/references/ref-en.txt

同一份指南对 retention_analysis 有一句关键限定:

Choose each relationship marker only within the reference role already defined for that label in subject_definitions. Do not treat newly added actions, backgrounds, or plot events in the target video as losses of reference fidelity.

出处:同上,§4 retention_analysis

这句话的直接推论:

你以为 fully_preserved 会锁住的官方实际锁住的
脸、发型、浴巾、身份是,如果这些写进了 subject 定义
她站在花洒下、双手抓浴巾、不走去开门否。 开门是「目标视频里新加的动作」,官方明确说这不算 retention 损失

所以「Subject 1 fully_preserved + 她全程站着不动」在官方语义里并不矛盾:模型可以完整保留她的脸和浴巾,同时让她去执行 prompt 里权重最高的动作。这正好对上你们 9 次失败的形态。

ComfyUI 官方 R2V 教程把同一件事说得更短:

Assign each reference a job: State which reference drives which part of the shot (identity, style, motion, camera, voice). Explicit assignments tend to work much better.

出处:https://docs.comfy.org/tutorials/video/minimax/minimax-h3

身份参考如果没有同时被分配「姿态/站位」这份工作,动作轴就会落到 prompt 文本上,而文本里最显著的动词是「握手—下压—推门」。

1.3 官方运镜词 POV = 「主体的视角」,不是「门外那只手」

基础 prompt guide 的运镜表把 POV 定义成:

Motion type POV — The subject's point of view

出处:https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/h3-prompt-writing/references/base-en.txt(§4.3) Hugging Face:https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md

中文社区直译也是「主体视角代入」:

出处:https://github.com/r600a-code/minimax-h3-prompt-skill/blob/main/SKILL.md

如果 <Subject 1> 是那个女生,再写 POV / 「第一人称 POV」,模型的合法解读是:从她的眼睛看出去。那只握门把手的手就变成她的手。这和「站在门外朝里看」是相反的镜头。

teskor-hub 的经验规则把 POV 用在「设备本身就是视点」时:

Never write camera as a noun the subject interacts with. H3 renders she holds the camera as a prop. … When the device is the viewpoint, use POV and never mention it at all; a visible outstretched arm is what sells the grip.

出处:https://github.com/teskor-hub/minimax-h3-skill/blob/main/SKILL.md(Hard rules) 同文 troubleshooting:https://github.com/teskor-hub/minimax-h3-skill/blob/main/references/troubleshooting.md(“A physical camera or phone appears in the shot”)

注意:teskor 这条是「主体自己举着设备看自己/看镜头」的自拍类镜头。你们的镜头是门外第三人的手推门,看向室内的 Subject 1。套用官方 POV 会把视点绑到 Subject 1 身上,正好强化「她来开门」。

1.4 社区已记录的相邻失败模式(同一机制,不同症状)

社区症状机制和本场景的对应
The subject performs the wrong activity词频而不是场景逻辑:反复出现的动作会赢「开门 / handle / push」在 prompt 里比「站在花洒下抓浴巾」更密、更具体
The subject rotates instead of the camera运镜是文本条件里最弱的轴,模型用主体动画顶替门在动 → 她的手臂去完成这个运动
A physical camera appears / raised hand comes up empty没有视觉 token 的物体不会稳定出现;未定义的「第二人」会被并进 Subject 1POV 手没有自己的 Subject / 参考图,只能借用女生的手
Negative prompt does nothing模板用 BasicGuider,CFG 实际为 1,没有负向通道「不能靠近门、不能碰门」即使改成正面句,只要开门动词还在,仍会执行
Voice/accent leaks between subjects主体局部条件会串ComfyUI #15454,声音版的「归属泄漏」
「身边有物」被理解成交互本仓库本地手册已踩过门把手在画面前景,最近的有手的人会被派去握它

原文与出处:

词频决定动作(empirical):

The subject performs the wrong activity — dancing instead of the described action. Cause. Word frequency, not scene logic. Count the cues in the description: if incidental movement is named five or six times … and the actual subject of the clip is named once near the end, the repeated activity is what the conditioning carries. … The model cannot count, so those phrases order nothing; the description reads as an unordered set of actions and the most frequent one wins.

出处:https://github.com/teskor-hub/minimax-h3-skill/blob/main/references/troubleshooting.md

运镜弱、主体动画顶替:

The subject rotates instead of the camera. Cause. Camera motion is the weakest axis in text conditioning; the model substitutes subject animation, which it knows far better.

出处:同上

没有参考的物体/第二人不会自己站住:

The subject raises a hand but the object never appears. Cause. Objects absent from the references rarely materialise mid-shot. Fix. Supply the object as its own <Subject N> with a source, or restructure so it is the viewpoint and therefore off-frame.

出处:同上

否定词无效:

The negative prompt does nothing. Cause. The ComfyUI template uses BasicGuider — one conditioning input, CFG effectively 1, no negative socket. Fix. Rewrite bans as positive statements of the desired state.

出处:同上

以及:Every mention, bans included, adds weight. … Name it once, with no prohibition attached.

出处:同文 “The event happens more times than asked”

主体条件串扰(声音,官方 issue):

MiniMax H3: multi-speaker voice/accent conditioning can leak between subjects in FL2VA and Ref2VA

出处:https://github.com/Comfy-Org/ComfyUI/issues/15454

Expected: voice characteristics specified for one subject should remain bound to that subject. Actual: voice conditioning appears to affect the generated audio globally.

这不是动作泄漏的直接证据,但是同一类「标签写了 Subject 1 / Subject 2,模型仍把条件混到默认主体上」。

Reddit:the viewer 会多出一个人,the camera 才保住朝向主体的视点:

Took a few little prompt adjustments here and there to get H3 to respect point-of-view. I found that if you refer to "the viewer" (ie, "she kicks the viewer"), H3 is more predisposed to include an actual second person. But if you refer to "the camera" (ie, "she kicks the camera"), it's more predisposed to keep the desired point-of-view perspective.

出处:https://www.reddit.com/r/StableDiffusion/comments/1w3k9ti/trying_out_a_consistent_pointofview_shot_with/

这条对你们是双刃剑:写 the viewer 可能真的长出第二个人(你们要的是一只手,不是一个完整路人);写 the camera 保住「朝向女生」的机位,但那只手仍然没有独立身份,动作会回到 Subject 1。正确做法是 机位用 camera 句,手用独立 Subject,而不是单独押其中一个词。

Virse 把 Ref2VA 的默认合同说死了:

A useful rule for H3 is: references define identity; prompts define change.

出处:https://www.virse.ai/blog/minimax-h3-character-consistency

参考图女生一旦被当成 identity,prompt 里的 change(开门)就由她来演。这是「参考角色被拉进镜头动作」的最短解释。

本仓库本地手册已记录的同类坑:

走两步就换动作:模型把「身边有物」理解成交互。写 Forbidden 清单,点名 00:02 不换动作。 若「人擦过镜头飞远、机位钉死」反复失败,优先改成 FL2VA……首尾帧比加否定句更能锁空间。

出处:本仓库 h3_prompt_guide.txt

门把手就是「身边有物」。前景有一个可交互的把手、画面里又只有一个有手的人,模型会完成这个交互。

1.5 三视图参考图本身也在帮倒忙

teskor 明确不推荐把 front/profile/back 拼成一张 grid 当主身份源:

A multi-view character sheet gives worse identity than plain photos. … The sheet is resized as a single image, so at ref_image_size: max a 3 × 3 grid gives roughly 680 px per panel against 2048 px for a dedicated still. … The grid's own structure — framed panels, gutters, a repeated figure, a seamless backdrop — is also visible content competing with the target scene.

出处:https://github.com/teskor-hub/minimax-h3-skill/blob/main/references/troubleshooting.md

三视图会让模型看到「同一个人的多个身体」,再叠上「开门」这个强动作,更容易把她的手复制到门框上。

1.6 机制小结(用来解释「为什么六段式正面表述仍失败」)

参考图女生 ──► 唯一有视觉 token 的人 ──► 默认施事者
prompt 里最密、最具体的动词 = 握把 / 下压 / 推门
fully_preserved 不锁站位(官方明文)
官方 POV = Subject 1 的眼睛(会把她变成开门的人)
CFG≈1,否定句没有通道
没有「镜头侧的手」的 Subject / 参考图 ──► 那只手只能从女生身上借
门把手在前景 ──► 「身边有物 = 交互」先验

所以这是 已知机制上的可预期失败,不是「再把六段式写得更长一点」就能好的问题。


2. 社区推荐解法清单(提示词 / 参数 / 工作流)

下面每条都带出处。按层分类,不按推荐顺序(推荐顺序见第 3 节)。

2.1 提示词层

A. 把镜头侧的手定义成独立 <Subject>,并尽量给它一张参考图

官方允许一个 Subject 是「物体 / 道具」,也允许一张图提供多个 Subject:

One subject may be defined by multiple reference assets, and one reference asset may provide multiple subjects.

出处:官方 ref guide,§2.1

teskor 对「手里突然要出现一个参考图里没有的东西」的修法就是给它自己的 Subject:

Supply the object as its own <Subject N> with a source

出处:https://github.com/teskor-hub/minimax-h3-skill/blob/main/references/troubleshooting.md

Kapwing 产品手部特写的推荐句式(把施事者写成「一只手从画框进入」,而不是「那个女人」):

[3–6s] A hand enters from frame right, grips the product securely, lifts it from the surface, and rotates it approximately 30 degrees until the front label faces camera.

出处:https://www.kapwing.com/resources/how-to-prompt-minimax-h3-hailuo-3-0-a-guide-for-ai-video-creators/

建议英文句式:

<Subject 2> is the camera-operator's right hand, a different person from <Subject 1>.
It belongs to the unseen adult standing in the corridor, on the camera side of the frosted door.
Only this hand and a short length of forearm are visible at the lower-left foreground edge.

再配一张「成人右手握金属门把手」的特写作为 <Picture 2>。没有这张图,Subject 2 仍然只是文本,模型还是会去借用 Subject 1 的手。

B. 把「站位 + 抓浴巾姿态」写进 Subject 定义,并在 retention 里当作 fully_preserved 的 role

官方:marker 只在「已定义角色」内生效。要把静止变成角色,而不是事后禁令。

ethanfel 的 Plan v2 把「动作 / 姿态」明确列为可绑定内容:

Use Subject Binding to: combine multiple references as evidence for one Subject; or transfer an attribute, pose, expression, or action to an existing Subject.

出处:https://github.com/ethanfel/ComfyUI-MiniMax-H3-Guide/blob/main/docs/PLAN_V2.md

官方 README 里就有用正面句子锁死双手的例子:

She stands perfectly still with her hands clasped tightly behind her back.

出处:https://github.com/MiniMax-AI/MiniMax-H3/blob/main/README.md(案例描述)

建议英文句式:

<Subject 1> is the young East Asian woman in <Picture 1>, wearing a light-blue bath towel.
Her standing pose is part of her defined role: she remains under the shower head in the far
depth of the bathroom, both hands holding the top edge of the towel at her chest, feet planted.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - face, hair, towel, planted stance under
the shower head, both hands on the towel, and her distance from the doorway stay as defined.

C. 机位用 the camera / Static Shot,不要用官方词 POV,也不要写 the viewer

建议英文句式:

The camera holds a static shot from the corridor, looking through the frosted-glass bathroom door
toward <Subject 1> in the far shower. The camera body stays outside the bathroom.

不要写:first-person POV、POV shot、the viewer pushes the door、she opens the door for the camera。

D. 手和机位拆成两句;手写解剖轨迹,只出现一次

teskor:camera 运动和身体动作必须分句,用所属动词连在一起会变成道具。Kapwing:短动词不够,要 starting state → movement → ending state。Hailuo 指南同样强调手、方向、注视:

Weak: “A woman opens a door dramatically.” Better: “A woman grips the metal handle with her right hand, pulls the heavy door toward her, steps backward, and looks through the opening.”

出处:https://www.videotoprompt.app/posts/hailuo-h3-prompt-guide

你们要把施事者从 woman 换成 camera-side hand:

<Subject 2>, a right hand at the lower-left foreground, wraps four fingers around the
horizontal metal handle. The thumb presses down, the wrist rotates, the latch releases,
and the frosted door swings inward toward the bathroom. Only this hand is on the handle.

整段 prompt 里「handle / push / open」只出现这一次。teskor:「Count the mentions in your own prompt; a failing prompt typically names the event six times。」

E. 词频倒置:Subject 1 的静止状态要比开门写得更密

teskor 的修法是把次要动作降成从句,把关键动作写成每个节拍自己的句子。对本场景要反过来:她的站位、双手、与门的距离在每个时间段重复;开门只在一个节拍写一次完整轨迹。

NikoDemon 补充:完全冻住会像卡死,给静止的人一点微动作:

Give the hold something to do. A held framing with nothing happening renders as a literal freeze, and two seconds of a motionless actor looks like the video stalled. Write in a breath, a weight shift, an eyeline change. The camera holds still, the performer doesn't.

出处:https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

建议: 蒸汽、水珠、一次眨眼、一次吸气。这些是「变化」,满足 teskor「Every clause should name something that changes」,又不会变成走去开门。

teskor 那条「She stands in the doorway is not an action」是在抱怨画面完全死掉,不是在鼓励人物去找门。用环境微动代替人物走位。

F. summary 里不要让 Subject 1 当开门的主语

官方:summary 用已定义标签描述主参考关系,不要在这里引入新标签。社区经验是 summary 会被当成整段的任务句。如果写成:

[reference generation] <Subject 1> in a bathroom as a door opens...

模型仍可能把开门派给她。应写成两个并列主语:

[reference generation] An 8-second single take from the corridor. <Subject 1> remains under
the shower in the far bathroom. <Subject 2> depresses the door handle in the foreground
and the frosted door swings inward, revealing <Subject 1> still in the same place.

G. 把动作从 summary 挪到 detailed_description 的时间线,并用序数而不是「for the first two seconds」

teskor:

Beats written as prose second-counts — for the first two and a half seconds … The model cannot count … Fix. Enumerate … joined by ordinals — she begins, next, then, finally

出处:troubleshooting.md

官方时间戳格式仍建议用 At 00:03.500 开后镜,但同一镜内用 begins / next / then / finally 比散文秒数更稳。

H. 参考职责写在 retention,而不是正文里禁

teskor / Kapwing 都强调:scope 放在 retention_analysis,正文里 do not copy 是在和视觉信号打架,且 CFG=1 时没有负向通道。

Put the scope in retention_analysis, where the format expects it: <Video 1> (camera movement and pacing): weak_reference - only the travelling path and handheld rhythm are followed; none of its people, wardrobe, location or lighting appear.

出处:troubleshooting.md

Kapwing:

Video 1 defines the hand trajectory and movement timing only. Do not copy its actor, wardrobe, setting, lighting, or camera movement.

出处:https://www.kapwing.com/resources/how-to-prompt-minimax-h3-hailuo-3-0-a-guide-for-ai-video-creators/

如果加一段「第一人称推门」参考视频,必须把视频角色的身份标成 weak_reference / attribute_transfer 到 <Subject 2>(那只手),绝不能 transfer 到 Subject 1。否则就是标准的 motion-transfer:女生会去跳那段推门。

I. 身份放进 <Subject N>,不要给三视图单独的 <Picture N> 行

官方:

If an image is used only to define a character, scene, costume, or style, do not create a standalone picture entry. Instead, cite the image source inside the corresponding <Subject N> definition.

出处:官方 ref guide §2.2

teskor 把这称为「the single most common structural error」。

J. 中文社区的「一张图只锁一件事」

APIDot 中文版:

关键规则:告诉 H3 每个参考具体控制什么。如果你不说,模型会尝试从每张图中使用所有信息。

出处:https://apidot.ai/zh/blog/minimax-h3-omni-reference-guide

xianyu110 提示词库:

参考图管身份,提示词管动作。

出处:https://github.com/xianyu110/awesome-minimax-h3-prompts/blob/main/README.zh-CN.md

这再次解释失败:你们把「身份」和「开门」写在同一条 prompt 里,又只给了女生的图,动作只能由她完成。解法不是再强调「她不做开门」,而是给开门另找一个有图的施事者。

K. 正面写「only this hand appears at the handle」,不要写「she does not touch the door」

teskor:否定会把被禁的事件加进 conditioning。Hailuo / Kapwing 也要求 constraints 用锁定元素而不是禁令。本仓库手册同样要求正面写死双手空摆。

建议:

Throughout the take, <Subject 1>'s hands stay on the towel at her chest.
The metal handle is gripped only by <Subject 2>.

2.2 参数层

参数建议原理与出处
ref_image_size迭代用 match,定稿用 maxmax 保身份明显更好,参考 token 每步都 attend,更慢。ComfyUI 官方 R2V tips:https://docs.comfy.org/tutorials/video/minimax/minimax-h3;teskor comfyui.md
denoise生成保持 1.0;只有 v2v / 二采 / inpaint 才降teskor:「Denoise 1.0. Only lower it for video-to-video.」
CFG / 负向不要指望 negative;不要轻易换 CFGGuiderBasicGuider,CFG≈1;换 CFGGuider 推理时间翻倍且训练行为未知
steps模板 20;turbo LoRA 才 4/8官方模板;提高步数对「谁去开门」几乎无帮助
scheduler在 simple 上 A/B beta / normalteskor:reference-heavy 图上这对是采样面板唯一值得试的
seed接近目标后立刻 fixedteskor:「Debugging a prompt against moving noise is guesswork.」
length和 prompt 时间线对齐;开门这类单动作 static 机位用 124(5.17 s)足够teskor:124 = 单动作 static;写 10 秒渲染 124 帧会把动作挤乱
encoder 精度身份走 VLM,Ref2VA 比 I2VA 更吃 encoderteskor:nvfp4_awq 可能弱化身份;有 VRAM 升 int8_convrot / bf16
turbo LoRA / cache排查动作归属时关掉社区普遍认为 skip-cache 伤角色一致性和动作;最终渲染关掉
ControlNet control_context_scale先 1.0 再往下调https://minimaxh3.cc/guides/minimax-h3-controlnet

没有在官方节点里找到名为「retention 强度」的独立滑条。retention 是 prompt 里的英文标记(fully_preserved 等),不是采样参数。部分第三方 Director(seesee75、muse-collective)把 retention 做成下拉,写进同一段 prompt,并不改变底层。

部分自定义节点暴露 ref_strength / ref_decay / ref_ramp(例如 https://github.com/kat3ri/ComfyUI-H3-Cast)。原生 MiniMaxH3ReferenceToVideo 文档里的对应旋钮是 ref_image_size。提高参考强度只能让女生更像参考图,不能阻止她去开门。

2.3 工作流层

L. 换 FL2VA:首尾帧把「她在深处、手在门上」钉死

官方 FL2VA:Picture 1 = 0.00 s,Picture 2 = 片尾,中间只写可观察的路径。本仓库手册:空间锁不住时,首尾帧比加否定句更有效。

teskor 的模式选择规则:

if the value is in the picture — its room, light, grain, composition — use fl2va. If the value is in who or what is in it, use ref2va. fl2va is animation: it takes the frame and moves it. ref2va is casting: it takes the subject and shoots a new scene.

出处:https://github.com/teskor-hub/minimax-h3-skill/blob/main/SKILL.md

你们真正要锁的是构图和站位(她在深处、手在前景门上),这是 picture 里的值,不是「再拍一条她的新戏」。FL2VA 更对症。

首帧:走廊看向磨砂门,门未开或微缝,深处花洒下女生抓浴巾,前景左下角一只与她无关的手握着把手。
尾帧:门已推开,女生仍在花洒下同一位置、同一抓巾姿势,前景的手在把门推开后的位置。

对齐句用官方固定字符串。checkpoint 换成 fl2va。

Runware / deAPI 也指出:I2VA/FL2VA 的首帧锁的是构图,新主体(「a hand entering from the right」)可以后加:

Your image fixes frame 0 and nothing else, so a hand entering from the right, a second person, a passing car are all fair game.

出处:https://deapi.ai/blog/minimax-h3-image-to-video-prompting-guide-alignment-lines-keyframes-and-3-example-prompts

M. MiniMaxH3AddGuide:在任意帧钉一张「她仍在花洒下」的静帧

ComfyUI 0.34+ / PR #15439 起,关键帧不再限于首尾:

This node anchors an image, a short clip, audio, or a clip with its soundtrack at any chosen frame.

出处:https://docs.comfy.org/built-in-nodes/MiniMaxH3AddGuide PR:https://github.com/Comfy-Org/ComfyUI/pull/15439 教程说明:https://docs.comfy.org/tutorials/video/minimax/minimax-h3(Anchoring guides at any frame)

做法:Ref2VA 或 FL2VA 出片的同时,把「女生在花洒下抓浴巾」的静帧(可从三视图裁正面)锚到 0、中间、末尾。门和前景手由 prompt 生成,女生的像素位置被 guide 拉住。

注意:teskor 写过本地 H3 从空 latent 起采样,keyframe/reference 是每步重注入、不被 denoise 的条件,不是逐帧复制。guide 是强条件,不是保证。

N. Fun ControlNet Pose:用一条「她站着抓浴巾」的 pose 视频冻住身体

sepiablue 验证过 Ref2VA + Fun ControlNet Union:参考图锁脸,pose 视频锁骨架。

做法:做一段 5 s 的 DWPose,人物站在画面深处、双手在胸前,几乎不动;门和前景手不要画进 pose(或单独用 depth/canny 管门)。H3 会让角色身体跟 pose,开门就很难再派给她。

O. 拆成两段生成再剪:先特写手推门,再看向室内的她

teskor / xianyu110 都强调 4–15 秒只压一个动作,多拍比一镜塞两件事稳。

  1. Clip A:近景,只有门把手和镜头侧的手,没有女生。T2VA 或 FL2VA 即可。
  2. Clip B:门已打开(或透过磨砂玻璃),女生在花洒下微动。Ref2VA 锁身份。
  3. 剪辑切在门开的瞬间。

这从结构上消灭「同一帧里两个施事者抢一只手」。

P. 掩膜 inpaint / 二采:先出「她站着」的片子,再只重绘门和前景手

mdkberry 记录了 drozbay 的 H3 masked inpainting(ComfyUI PR #15375,节点 https://github.com/drozbay/MaskVidExperiments):

This is really useful for swapping faces at distance or swapping characters in to a video.

出处:https://github.com/mdkberry/comfyui_workflows/tree/main/workflows_by_model/Minimax-H3

Fun ControlNet Union 也带 inpainting:白蒙版重绘、黑区保留。流程:

  1. Ref2VA 生成女生在浴室站着(prompt 里根本不提门把手)。
  2. 把前景门+把手区域涂白,二采 denoise 0.3–0.5,只让那一块长出「镜头侧的手在推门」。

Q. 不要用「角色替换 / video editing」把女生替换进一段推门视频

角色替换的标准合同是:

image = identity, video = performance

出处:https://www.virse.ai/blog/how-to-replace-a-character-in-video-with-minimax-h3

如果参考视频是「一个人在开门」,identity 换成你们的女生,结果一定是她去开门。这是本问题的反向操作。只有当参考视频是第一人称、开门的人永远只有一截前臂、室内深处另有一个站着的人时,才考虑 video 参考。

R. 参考视频只借「手部轨迹」,人必须不在画面里

teskor:

If you want camera motion, shoot the reference orbiting an object with no person in frame. Otherwise crop the reference so the face never appears.

出处:troubleshooting.md / comfyui.md 末尾

Kapwing / flaqai:Video 1 controls hand movement timing only.

自制 3–5 秒第一人称握把下压视频(不要拍到脸和身体),Ref2VA 里:

<Subject 2> is the camera-side hand whose gripping motion comes from <Video 1>.
<Video 1> (hand trajectory only): attribute_transfer - only the handle grip and downward press
transfer onto <Subject 2>; none of its people or room appear.

S. 三视图拆开,或降级为弱参考

见 §1.5。至少再加一张单独的正面半身,作为 Subject 1 的主身份;三视图整图标 weak_reference 或只当服装/浴巾形状。


3. 针对本场景最可能有效的 Top 3(按可行性)

可行性 = 改动小、和 Dasiwa REF2VA 现有图兼容、对「她去开门」这个具体失败的针对性强。不是「理论上最强」。

方案 1(先做):双 Subject + 手部参考图 + 禁用官方 POV 词 + 词频倒置

改什么: 仍走 Ref2VA / Dasiwa Director 六段式。新增一张「成人右手握金属门把手」特写接到 ref_image_1。重写 subject / summary / retention / detailed_description。不要改 checkpoint。

为什么最可能立刻见效:

  1. 失败根因是「画面里只有一个有手的人」。给开门的手独立视觉 token,是 teskor 对「raised hand comes up empty」的官方兼容修法,也是 Kapwing「A hand enters from frame right」的句式。
  2. 官方 fully_preserved 只保已定义角色——把「花洒下抓浴巾」写进 Subject 1 的 role,静止才进入 retention 合同。
  3. 不用 POV / the viewer,避免视点绑到 Subject 1 或长出完整第二人。
  4. 开门轨迹只写一次,她的站位每个节拍都写,对冲词频先验。
  5. 不换模式、不换工作流,和现有 9 次迭代的差别是结构而不是「再写一遍不要开门」。

原理一句话: 把「谁有手」从文本协商改成参考图协商。H3 对视觉 token 的服从远强于对禁令的服从。

可粘贴骨架:

subject_definitions:
<Subject 1> is the young East Asian woman whose appearance comes from <Picture 1>: fair skin,
dark hair, a light-blue bath towel wrapped at the chest. Her defined role includes her planted
stance: she occupies the far shower under the shower head, both hands holding the top edge of
the towel at her chest, body facing the doorway but remaining in the shower bay.
<Subject 2> is the camera-operator's right hand in <Picture 2>, a different adult from
<Subject 1>. Only this hand and a short forearm occupy the lower-left foreground, on the
corridor side of the frosted-glass door, wrapping the metal handle.
<Subject 3> is the bathroom interior: frosted-glass door in the near plane, shower fixtures
and steam in the far plane.

summary:
[reference generation] An 8-second single continuous take from the corridor. <Subject 1>
remains under the shower in the far bathroom, hands on the towel. <Subject 2> depresses the
handle and the frosted door swings inward, revealing <Subject 1> still in the same place.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - face, hair, towel, planted stance under
the shower head, both hands on the towel, and her distance from the doorway stay as defined.
<Subject 2> (appears in [Shot 1]): fully_preserved - the camera-side right hand on the handle;
it is the only hand that contacts the door.
<Subject 3> (appears in [Shot 1]): partially_preserved - bathroom layout and frosted glass stay;
the door leaf swings inward.

detailed_description:
The target video is live-action, photoreal, humid bathroom light, vertical 9:16.
[Shot 1] The camera holds a static shot from the corridor, looking through the frosted-glass
door toward <Subject 1> in the far shower. Steam drifts. Water beads on the glass.
Begins: <Subject 1> stands under the shower head, both hands holding the towel at her chest;
she blinks once; her feet stay planted. <Subject 2> is already on the handle at the lower-left
foreground.
Next: <Subject 2> wraps the metal handle, the thumb presses down, the latch releases, and the
frosted door swings inward. Only <Subject 2> is on the handle.
Then: the opening glass reveals more of the shower. <Subject 1> is still under the shower head,
hands still on the towel, the same distance from the doorway. She draws one breath. Water
continues to run. The camera does not move.
Finally: the door stays ajar. <Subject 1> holds the same stance. <Subject 2> rests on the
handle.

overall_soundscape:
Bathroom extractor hum and shower spray continue. A short metallic latch click, then the soft
sweep of frosted glass. Light fabric tension from the towel. One quiet breath.

non_diegetic_music:
N/A

配套参数: ref_image_size=max(定稿)、seed 固定、length 124 或 158、关掉 turbo/cache、denoise 1.0。三视图若必须保留,加一张单独正面半身当 Picture 1,三视图整图降为弱参考。

风险: 没有手部参考图时,Subject 2 仍可能被画成她的手。手入画且朝镜头时 teskor 警告多指,写 five clearly separated fingers, a firm grip,并让手停在画面下缘,不要伸到画面中心。

方案 2(方案 1 仍抢手时):FL2VA 首尾帧锁空间

改什么: Dasiwa 切到 I2VA/FL2VA(fl2va 权重)。用图像工具或一次静帧生成做出两张构图图(不必是真实摄影):

为什么排第二:

  1. 官方和本仓库都认为空间问题用关键帧比用禁令强。你们要锁的正是「她不要走到门边」。
  2. teskor 的模式分叉:构图在图里 → fl2va。
  3. 两端都画出「她在深处 + 手在门上」,中间插值很难再让她走过来——走过来会违反尾帧。
  4. 改动中等:要换权重、做两张图,但不用 ControlNet。

原理一句话: 把「她始终在深处」从文本条件升级为每步重注入的像素条件。

提示词要点: 官方对齐句必须是第一行;正文只写门如何从首帧摆到尾帧;写 settles into the pose, spacing, and composition established by Picture 2。不要写 POV。

可叠加 MiniMaxH3AddGuide,把同一张「她在花洒下」的静帧锚到中间帧,进一步防中段走位。

风险: FL2VA 不吃角色三视图那套 <Subject> 标签,身份一致性可能略差;两张构图图的光线/颗粒必须接近,否则会变成「美颜 morph」(teskor templates.md)。门的运动要是两帧之间能插值的,不要首帧关门、尾帧人已经换房间。

方案 3(要稳、可加管线时):Pose ControlNet 冻身体,或拆镜 / 掩膜二采

改什么(三选一,强度递增):

  1. 拆镜: 5 s 手+门特写(无人)+ 5 s 门开后看她(Ref2VA)。剪辑解决归属。
  2. Pose ControlNet: 一条「双手抓巾、站在深处」的几乎静止 pose 视频驱动 Fun ControlNet Union,Ref2VA 仍喂女生图。ComfyUI 2026-09-01 已允许 reference + Fun Union 共存。
  3. 先出站立片再 inpaint 门和手: drozbay MaskVid / Fun inpaint,denoise 0.3–0.5 只动前景。

为什么排第三: 对「身体不要去开门」几乎是硬约束,但 Dasiwa 默认 REF2VA 图要接新节点、新模型(minimax_h3_fun_controlnet_union_*.safetensors)或二采图,排查成本最高。

原理一句话: 骨架/掩膜是空间约束,不走「文本里谁是主语」这条会串的通道。

风险: pose 视频如果把手也画成去握门,问题会原样回来——pose 里她的手必须一直在胸口。拆镜会失去「一镜从门外推开看到她」的连贯感,需要剪辑补。


4. 明确不建议的做法

做法原因
继续加长六段式、反复写「她不做开门」CFG=1,禁令加权重;你们已经 9 次
使用官方运镜词 POV / first-person POV官方定义为 subject's point of view
写 the viewer opens the doorReddit:容易长出完整第二人
video editing 把女生替换进「有人开门」的视频image=identity, video=performance,她会去演开门
指望 ref_strength / 更高 steps / 换 seed 解决归属这些不改变默认施事者
只丢三视图、不给手部图grid 还在和场景抢内容,且开门的手没有视觉源
把开门动作在 summary、retention、detailed_description 各写一遍词频让开门赢

5. 建议的试验顺序(一次只改一维)

官方社区(minimaxh3.studio / teskor)都要求每次只改一个控制变量。

  1. 同一 seed,只改 prompt 为方案 1 骨架,仍不接手部图。 看她是否还去握门。若手变成「悬浮的第二只手」但属于她,说明还缺视觉源。
  2. 同一 prompt,加上手部特写 Picture 2,ref_image_size=max。 这是方案 1 的完整形态。
  3. 若她仍走位:做首尾帧,切 FL2VA(方案 2),prompt 几乎只描述门的运动。
  4. 若身份漂了但站位对了:FL2VA 出片后再用 Ref2VA 低 denoise + 参考图修脸(mdkberry detailer 路)。
  5. 若必须一镜且手和人都要稳:Pose ControlNet 或 inpaint(方案 3)。

6. 主要出处

官方

ComfyUI

社区技能与手册

英文社区

中文社区

工作流 / ControlNet


7. 一句话给后续迭代用

Ref2VA 不会「理解」第一人称推门 vs 室内站着的人;它会把参考图里那个有手的人,派去完成 prompt 里最具体的那组手部动作。要分开这两件事,必须让开门的手拥有自己的 Subject 和自己的像素,或者用首尾帧 / pose / 掩膜在像素层锁住女生的身体。继续在六段式里用中文或英文强调「她不能开门」,官方 retention 语义和 CFG=1 都不会帮你执行这条禁令。

下载此文件