场景: ComfyUI Dasiwa 工作流,REF2VA(参考图生视频)。参考图是浅蓝浴巾亚洲女生三视图。提示词要求第一人称站在门外,镜头侧一只手握住门把手下压推开磨砂玻璃门;参考图女生(<Subject 1>)全程站在房间深处花洒下方,双手抓住胸前浴巾顶部,不靠近门、不碰门、不做开门动作。
实测失败: 9 个版本迭代后,Subject 1 仍「同步执行开门动作」——门开时她的手出现在门框/门把手上,或她走到门边扶门。开门被模型当成她在做,而不是 POV 镜头侧那只手。
已试过且仍失败: 正面表述(无否定词)、官方六段式(subject_definitions / summary / retention_analysis / detailed_description / overall_soundscape / non_diegetic_music)。
调研范围: MiniMax 官方 prompt guide、ComfyUI 原生文档与 GitHub issues、ethanfel/ComfyUI-MiniMax-H3-Guide、teskor-hub/minimax-h3-skill、Reddit r/StableDiffusion、Kapwing / APIDot / Hailuo 中文指南、Fun ControlNet / Add Guide 工作流资料。截止日期约 2026-09-02。
结论先行: 这不是六段式写错那么简单。它是 Ref2VA 的结构性偏向:参考图人物是画面里唯一有视觉 token 的「施事者」,提示词里任何显著肢体动作(握手、下压、推门)都会默认落到这个人身上。官方 fully_preserved 不锁姿态/站位,官方运镜词 POV 的字面意思是「主体视角」而不是「门外第三人称手」。社区没有一篇 issue 用完全相同的标题描述「参考角色被拉进镜头/第二人的动作」,但同一机制在多处被反复踩到,解法也已经收敛。
在 Comfy-Org/ComfyUI issues、MiniMax 官方 guide、ethanfel 仓库里,没有找到一条标题等价于:
reference character gets pulled into actions described for the camera / POV / second person
所以不宜把它说成「官方已确认的 named bug」。更准确的说法是:这是 Ref2VA 默认施事者绑定 + 官方 retention 语义 + 官方 POV 词义,三者叠加后的可复现失败模式。 同类现象在社区已经被写成可操作的 troubleshooting。
fully_preserved 保的是「已定义角色」,不是「站着别动」MiniMax 全参考模式指南把 <Subject N> 定义为「会实际出现在目标视频里的可复用可见内容」,范围包括人、物、场景、服装,也包括动作、表情、姿态:
<Subject N>is used for reusable visible content, including: People, animals, or objects; Scenes, backgrounds, or environments; Clothing, props, interfaces, or visual effects; Styles, actions, expressions, or poses.出处:https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md 镜像:https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/h3-prompt-writing/references/ref-en.txt
同一份指南对 retention_analysis 有一句关键限定:
Choose each relationship marker only within the reference role already defined for that label in
subject_definitions. Do not treat newly added actions, backgrounds, or plot events in the target video as losses of reference fidelity.出处:同上,§4
retention_analysis
这句话的直接推论:
你以为 fully_preserved 会锁住的 | 官方实际锁住的 |
|---|---|
| 脸、发型、浴巾、身份 | 是,如果这些写进了 subject 定义 |
| 她站在花洒下、双手抓浴巾、不走去开门 | 否。 开门是「目标视频里新加的动作」,官方明确说这不算 retention 损失 |
所以「Subject 1 fully_preserved + 她全程站着不动」在官方语义里并不矛盾:模型可以完整保留她的脸和浴巾,同时让她去执行 prompt 里权重最高的动作。这正好对上你们 9 次失败的形态。
ComfyUI 官方 R2V 教程把同一件事说得更短:
Assign each reference a job: State which reference drives which part of the shot (identity, style, motion, camera, voice). Explicit assignments tend to work much better.
出处:https://docs.comfy.org/tutorials/video/minimax/minimax-h3
身份参考如果没有同时被分配「姿态/站位」这份工作,动作轴就会落到 prompt 文本上,而文本里最显著的动词是「握手—下压—推门」。
POV = 「主体的视角」,不是「门外那只手」基础 prompt guide 的运镜表把 POV 定义成:
Motion type
POV— The subject's point of view出处:https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/h3-prompt-writing/references/base-en.txt(§4.3) Hugging Face:https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
中文社区直译也是「主体视角代入」:
出处:https://github.com/r600a-code/minimax-h3-prompt-skill/blob/main/SKILL.md
如果 <Subject 1> 是那个女生,再写 POV / 「第一人称 POV」,模型的合法解读是:从她的眼睛看出去。那只握门把手的手就变成她的手。这和「站在门外朝里看」是相反的镜头。
teskor-hub 的经验规则把 POV 用在「设备本身就是视点」时:
Never write camera as a noun the subject interacts with. H3 renders
she holds the cameraas a prop. … When the device is the viewpoint, usePOVand never mention it at all; a visible outstretched arm is what sells the grip.出处:https://github.com/teskor-hub/minimax-h3-skill/blob/main/SKILL.md(Hard rules) 同文 troubleshooting:https://github.com/teskor-hub/minimax-h3-skill/blob/main/references/troubleshooting.md(“A physical camera or phone appears in the shot”)
注意:teskor 这条是「主体自己举着设备看自己/看镜头」的自拍类镜头。你们的镜头是门外第三人的手推门,看向室内的 Subject 1。套用官方 POV 会把视点绑到 Subject 1 身上,正好强化「她来开门」。
| 社区症状 | 机制 | 和本场景的对应 |
|---|---|---|
| The subject performs the wrong activity | 词频而不是场景逻辑:反复出现的动作会赢 | 「开门 / handle / push」在 prompt 里比「站在花洒下抓浴巾」更密、更具体 |
| The subject rotates instead of the camera | 运镜是文本条件里最弱的轴,模型用主体动画顶替 | 门在动 → 她的手臂去完成这个运动 |
| A physical camera appears / raised hand comes up empty | 没有视觉 token 的物体不会稳定出现;未定义的「第二人」会被并进 Subject 1 | POV 手没有自己的 Subject / 参考图,只能借用女生的手 |
| Negative prompt does nothing | 模板用 BasicGuider,CFG 实际为 1,没有负向通道 | 「不能靠近门、不能碰门」即使改成正面句,只要开门动词还在,仍会执行 |
| Voice/accent leaks between subjects | 主体局部条件会串 | ComfyUI #15454,声音版的「归属泄漏」 |
| 「身边有物」被理解成交互 | 本仓库本地手册已踩过 | 门把手在画面前景,最近的有手的人会被派去握它 |
原文与出处:
词频决定动作(empirical):
The subject performs the wrong activity — dancing instead of the described action. Cause. Word frequency, not scene logic. Count the cues in the description: if incidental movement is named five or six times … and the actual subject of the clip is named once near the end, the repeated activity is what the conditioning carries. … The model cannot count, so those phrases order nothing; the description reads as an unordered set of actions and the most frequent one wins.
出处:https://github.com/teskor-hub/minimax-h3-skill/blob/main/references/troubleshooting.md
运镜弱、主体动画顶替:
The subject rotates instead of the camera. Cause. Camera motion is the weakest axis in text conditioning; the model substitutes subject animation, which it knows far better.
出处:同上
没有参考的物体/第二人不会自己站住:
The subject raises a hand but the object never appears. Cause. Objects absent from the references rarely materialise mid-shot. Fix. Supply the object as its own
<Subject N>with a source, or restructure so it is the viewpoint and therefore off-frame.出处:同上
否定词无效:
The negative prompt does nothing. Cause. The ComfyUI template uses
BasicGuider— one conditioning input, CFG effectively 1, no negative socket. Fix. Rewrite bans as positive statements of the desired state.出处:同上
以及:Every mention, bans included, adds weight. … Name it once, with no prohibition attached.
出处:同文 “The event happens more times than asked”
主体条件串扰(声音,官方 issue):
MiniMax H3: multi-speaker voice/accent conditioning can leak between subjects in FL2VA and Ref2VA
出处:https://github.com/Comfy-Org/ComfyUI/issues/15454
Expected: voice characteristics specified for one subject should remain bound to that subject. Actual: voice conditioning appears to affect the generated audio globally.
这不是动作泄漏的直接证据,但是同一类「标签写了 Subject 1 / Subject 2,模型仍把条件混到默认主体上」。
Reddit:the viewer 会多出一个人,the camera 才保住朝向主体的视点:
Took a few little prompt adjustments here and there to get H3 to respect point-of-view. I found that if you refer to "the viewer" (ie, "she kicks the viewer"), H3 is more predisposed to include an actual second person. But if you refer to "the camera" (ie, "she kicks the camera"), it's more predisposed to keep the desired point-of-view perspective.
这条对你们是双刃剑:写 the viewer 可能真的长出第二个人(你们要的是一只手,不是一个完整路人);写 the camera 保住「朝向女生」的机位,但那只手仍然没有独立身份,动作会回到 Subject 1。正确做法是 机位用 camera 句,手用独立 Subject,而不是单独押其中一个词。
Virse 把 Ref2VA 的默认合同说死了:
A useful rule for H3 is: references define identity; prompts define change.
出处:https://www.virse.ai/blog/minimax-h3-character-consistency
参考图女生一旦被当成 identity,prompt 里的 change(开门)就由她来演。这是「参考角色被拉进镜头动作」的最短解释。
本仓库本地手册已记录的同类坑:
走两步就换动作:模型把「身边有物」理解成交互。写 Forbidden 清单,点名 00:02 不换动作。 若「人擦过镜头飞远、机位钉死」反复失败,优先改成 FL2VA……首尾帧比加否定句更能锁空间。
出处:本仓库
h3_prompt_guide.txt
门把手就是「身边有物」。前景有一个可交互的把手、画面里又只有一个有手的人,模型会完成这个交互。
teskor 明确不推荐把 front/profile/back 拼成一张 grid 当主身份源:
A multi-view character sheet gives worse identity than plain photos. … The sheet is resized as a single image, so at
ref_image_size: maxa 3 × 3 grid gives roughly 680 px per panel against 2048 px for a dedicated still. … The grid's own structure — framed panels, gutters, a repeated figure, a seamless backdrop — is also visible content competing with the target scene.出处:https://github.com/teskor-hub/minimax-h3-skill/blob/main/references/troubleshooting.md
三视图会让模型看到「同一个人的多个身体」,再叠上「开门」这个强动作,更容易把她的手复制到门框上。
参考图女生 ──► 唯一有视觉 token 的人 ──► 默认施事者
prompt 里最密、最具体的动词 = 握把 / 下压 / 推门
fully_preserved 不锁站位(官方明文)
官方 POV = Subject 1 的眼睛(会把她变成开门的人)
CFG≈1,否定句没有通道
没有「镜头侧的手」的 Subject / 参考图 ──► 那只手只能从女生身上借
门把手在前景 ──► 「身边有物 = 交互」先验
所以这是 已知机制上的可预期失败,不是「再把六段式写得更长一点」就能好的问题。
下面每条都带出处。按层分类,不按推荐顺序(推荐顺序见第 3 节)。
<Subject>,并尽量给它一张参考图官方允许一个 Subject 是「物体 / 道具」,也允许一张图提供多个 Subject:
One subject may be defined by multiple reference assets, and one reference asset may provide multiple subjects.
出处:官方 ref guide,§2.1
teskor 对「手里突然要出现一个参考图里没有的东西」的修法就是给它自己的 Subject:
Supply the object as its own
<Subject N>with a source出处:https://github.com/teskor-hub/minimax-h3-skill/blob/main/references/troubleshooting.md
Kapwing 产品手部特写的推荐句式(把施事者写成「一只手从画框进入」,而不是「那个女人」):
[3–6s] A hand enters from frame right, grips the product securely, lifts it from the surface, and rotates it approximately 30 degrees until the front label faces camera.出处:https://www.kapwing.com/resources/how-to-prompt-minimax-h3-hailuo-3-0-a-guide-for-ai-video-creators/
建议英文句式:
<Subject 2> is the camera-operator's right hand, a different person from <Subject 1>.
It belongs to the unseen adult standing in the corridor, on the camera side of the frosted door.
Only this hand and a short length of forearm are visible at the lower-left foreground edge.
再配一张「成人右手握金属门把手」的特写作为 <Picture 2>。没有这张图,Subject 2 仍然只是文本,模型还是会去借用 Subject 1 的手。
官方:marker 只在「已定义角色」内生效。要把静止变成角色,而不是事后禁令。
ethanfel 的 Plan v2 把「动作 / 姿态」明确列为可绑定内容:
Use Subject Binding to: combine multiple references as evidence for one Subject; or transfer an attribute, pose, expression, or action to an existing Subject.
出处:https://github.com/ethanfel/ComfyUI-MiniMax-H3-Guide/blob/main/docs/PLAN_V2.md
官方 README 里就有用正面句子锁死双手的例子:
She stands perfectly still with her hands clasped tightly behind her back.
出处:https://github.com/MiniMax-AI/MiniMax-H3/blob/main/README.md(案例描述)
建议英文句式:
<Subject 1> is the young East Asian woman in <Picture 1>, wearing a light-blue bath towel.
Her standing pose is part of her defined role: she remains under the shower head in the far
depth of the bathroom, both hands holding the top edge of the towel at her chest, feet planted.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - face, hair, towel, planted stance under
the shower head, both hands on the towel, and her distance from the doorway stay as defined.
the camera / Static Shot,不要用官方词 POV,也不要写 the viewerPOV = subject's point of view(见 §1.3)the viewer 容易长出第二个人;the camera 保住朝向主体的视点(见 §1.4)建议英文句式:
The camera holds a static shot from the corridor, looking through the frosted-glass bathroom door
toward <Subject 1> in the far shower. The camera body stays outside the bathroom.
不要写:first-person POV、POV shot、the viewer pushes the door、she opens the door for the camera。
teskor:camera 运动和身体动作必须分句,用所属动词连在一起会变成道具。Kapwing:短动词不够,要 starting state → movement → ending state。Hailuo 指南同样强调手、方向、注视:
Weak: “A woman opens a door dramatically.” Better: “A woman grips the metal handle with her right hand, pulls the heavy door toward her, steps backward, and looks through the opening.”
出处:https://www.videotoprompt.app/posts/hailuo-h3-prompt-guide
你们要把施事者从 woman 换成 camera-side hand:
<Subject 2>, a right hand at the lower-left foreground, wraps four fingers around the
horizontal metal handle. The thumb presses down, the wrist rotates, the latch releases,
and the frosted door swings inward toward the bathroom. Only this hand is on the handle.
整段 prompt 里「handle / push / open」只出现这一次。teskor:「Count the mentions in your own prompt; a failing prompt typically names the event six times。」
teskor 的修法是把次要动作降成从句,把关键动作写成每个节拍自己的句子。对本场景要反过来:她的站位、双手、与门的距离在每个时间段重复;开门只在一个节拍写一次完整轨迹。
NikoDemon 补充:完全冻住会像卡死,给静止的人一点微动作:
Give the hold something to do. A held framing with nothing happening renders as a literal freeze, and two seconds of a motionless actor looks like the video stalled. Write in a breath, a weight shift, an eyeline change. The camera holds still, the performer doesn't.
建议: 蒸汽、水珠、一次眨眼、一次吸气。这些是「变化」,满足 teskor「Every clause should name something that changes」,又不会变成走去开门。
teskor 那条「She stands in the doorway is not an action」是在抱怨画面完全死掉,不是在鼓励人物去找门。用环境微动代替人物走位。
官方:summary 用已定义标签描述主参考关系,不要在这里引入新标签。社区经验是 summary 会被当成整段的任务句。如果写成:
[reference generation] <Subject 1> in a bathroom as a door opens...
模型仍可能把开门派给她。应写成两个并列主语:
[reference generation] An 8-second single take from the corridor. <Subject 1> remains under
the shower in the far bathroom. <Subject 2> depresses the door handle in the foreground
and the frosted door swings inward, revealing <Subject 1> still in the same place.
teskor:
Beats written as prose second-counts —
for the first two and a half seconds… The model cannot count … Fix. Enumerate … joined by ordinals — she begins, next, then, finally出处:troubleshooting.md
官方时间戳格式仍建议用 At 00:03.500 开后镜,但同一镜内用 begins / next / then / finally 比散文秒数更稳。
teskor / Kapwing 都强调:scope 放在 retention_analysis,正文里 do not copy 是在和视觉信号打架,且 CFG=1 时没有负向通道。
Put the scope in
retention_analysis, where the format expects it:<Video 1> (camera movement and pacing): weak_reference - only the travelling path and handheld rhythm are followed; none of its people, wardrobe, location or lighting appear.出处:troubleshooting.md
Kapwing:
Video 1 defines the hand trajectory and movement timing only. Do not copy its actor, wardrobe, setting, lighting, or camera movement.
出处:https://www.kapwing.com/resources/how-to-prompt-minimax-h3-hailuo-3-0-a-guide-for-ai-video-creators/
如果加一段「第一人称推门」参考视频,必须把视频角色的身份标成 weak_reference / attribute_transfer 到 <Subject 2>(那只手),绝不能 transfer 到 Subject 1。否则就是标准的 motion-transfer:女生会去跳那段推门。
<Subject N>,不要给三视图单独的 <Picture N> 行官方:
If an image is used only to define a character, scene, costume, or style, do not create a standalone picture entry. Instead, cite the image source inside the corresponding
<Subject N>definition.出处:官方 ref guide §2.2
teskor 把这称为「the single most common structural error」。
APIDot 中文版:
关键规则:告诉 H3 每个参考具体控制什么。如果你不说,模型会尝试从每张图中使用所有信息。
出处:https://apidot.ai/zh/blog/minimax-h3-omni-reference-guide
xianyu110 提示词库:
参考图管身份,提示词管动作。
出处:https://github.com/xianyu110/awesome-minimax-h3-prompts/blob/main/README.zh-CN.md
这再次解释失败:你们把「身份」和「开门」写在同一条 prompt 里,又只给了女生的图,动作只能由她完成。解法不是再强调「她不做开门」,而是给开门另找一个有图的施事者。
teskor:否定会把被禁的事件加进 conditioning。Hailuo / Kapwing 也要求 constraints 用锁定元素而不是禁令。本仓库手册同样要求正面写死双手空摆。
建议:
Throughout the take, <Subject 1>'s hands stay on the towel at her chest.
The metal handle is gripped only by <Subject 2>.
| 参数 | 建议 | 原理与出处 |
|---|---|---|
ref_image_size | 迭代用 match,定稿用 max | max 保身份明显更好,参考 token 每步都 attend,更慢。ComfyUI 官方 R2V tips:https://docs.comfy.org/tutorials/video/minimax/minimax-h3;teskor comfyui.md |
| denoise | 生成保持 1.0;只有 v2v / 二采 / inpaint 才降 | teskor:「Denoise 1.0. Only lower it for video-to-video.」 |
| CFG / 负向 | 不要指望 negative;不要轻易换 CFGGuider | BasicGuider,CFG≈1;换 CFGGuider 推理时间翻倍且训练行为未知 |
| steps | 模板 20;turbo LoRA 才 4/8 | 官方模板;提高步数对「谁去开门」几乎无帮助 |
| scheduler | 在 simple 上 A/B beta / normal | teskor:reference-heavy 图上这对是采样面板唯一值得试的 |
| seed | 接近目标后立刻 fixed | teskor:「Debugging a prompt against moving noise is guesswork.」 |
| length | 和 prompt 时间线对齐;开门这类单动作 static 机位用 124(5.17 s)足够 | teskor:124 = 单动作 static;写 10 秒渲染 124 帧会把动作挤乱 |
| encoder 精度 | 身份走 VLM,Ref2VA 比 I2VA 更吃 encoder | teskor:nvfp4_awq 可能弱化身份;有 VRAM 升 int8_convrot / bf16 |
| turbo LoRA / cache | 排查动作归属时关掉 | 社区普遍认为 skip-cache 伤角色一致性和动作;最终渲染关掉 |
ControlNet control_context_scale | 先 1.0 再往下调 | https://minimaxh3.cc/guides/minimax-h3-controlnet |
没有在官方节点里找到名为「retention 强度」的独立滑条。retention 是 prompt 里的英文标记(fully_preserved 等),不是采样参数。部分第三方 Director(seesee75、muse-collective)把 retention 做成下拉,写进同一段 prompt,并不改变底层。
部分自定义节点暴露 ref_strength / ref_decay / ref_ramp(例如 https://github.com/kat3ri/ComfyUI-H3-Cast)。原生 MiniMaxH3ReferenceToVideo 文档里的对应旋钮是 ref_image_size。提高参考强度只能让女生更像参考图,不能阻止她去开门。
官方 FL2VA:Picture 1 = 0.00 s,Picture 2 = 片尾,中间只写可观察的路径。本仓库手册:空间锁不住时,首尾帧比加否定句更有效。
teskor 的模式选择规则:
if the value is in the picture — its room, light, grain, composition — use
fl2va. If the value is in who or what is in it, useref2va.fl2vais animation: it takes the frame and moves it.ref2vais casting: it takes the subject and shoots a new scene.出处:https://github.com/teskor-hub/minimax-h3-skill/blob/main/SKILL.md
你们真正要锁的是构图和站位(她在深处、手在前景门上),这是 picture 里的值,不是「再拍一条她的新戏」。FL2VA 更对症。
首帧:走廊看向磨砂门,门未开或微缝,深处花洒下女生抓浴巾,前景左下角一只与她无关的手握着把手。
尾帧:门已推开,女生仍在花洒下同一位置、同一抓巾姿势,前景的手在把门推开后的位置。
对齐句用官方固定字符串。checkpoint 换成 fl2va。
Runware / deAPI 也指出:I2VA/FL2VA 的首帧锁的是构图,新主体(「a hand entering from the right」)可以后加:
Your image fixes frame 0 and nothing else, so a hand entering from the right, a second person, a passing car are all fair game.
MiniMaxH3AddGuide:在任意帧钉一张「她仍在花洒下」的静帧ComfyUI 0.34+ / PR #15439 起,关键帧不再限于首尾:
This node anchors an image, a short clip, audio, or a clip with its soundtrack at any chosen frame.
出处:https://docs.comfy.org/built-in-nodes/MiniMaxH3AddGuide PR:https://github.com/Comfy-Org/ComfyUI/pull/15439 教程说明:https://docs.comfy.org/tutorials/video/minimax/minimax-h3(Anchoring guides at any frame)
做法:Ref2VA 或 FL2VA 出片的同时,把「女生在花洒下抓浴巾」的静帧(可从三视图裁正面)锚到 0、中间、末尾。门和前景手由 prompt 生成,女生的像素位置被 guide 拉住。
注意:teskor 写过本地 H3 从空 latent 起采样,keyframe/reference 是每步重注入、不被 denoise 的条件,不是逐帧复制。guide 是强条件,不是保证。
sepiablue 验证过 Ref2VA + Fun ControlNet Union:参考图锁脸,pose 视频锁骨架。
做法:做一段 5 s 的 DWPose,人物站在画面深处、双手在胸前,几乎不动;门和前景手不要画进 pose(或单独用 depth/canny 管门)。H3 会让角色身体跟 pose,开门就很难再派给她。
teskor / xianyu110 都强调 4–15 秒只压一个动作,多拍比一镜塞两件事稳。
这从结构上消灭「同一帧里两个施事者抢一只手」。
mdkberry 记录了 drozbay 的 H3 masked inpainting(ComfyUI PR #15375,节点 https://github.com/drozbay/MaskVidExperiments):
This is really useful for swapping faces at distance or swapping characters in to a video.
出处:https://github.com/mdkberry/comfyui_workflows/tree/main/workflows_by_model/Minimax-H3
Fun ControlNet Union 也带 inpainting:白蒙版重绘、黑区保留。流程:
角色替换的标准合同是:
image = identity, video = performance
出处:https://www.virse.ai/blog/how-to-replace-a-character-in-video-with-minimax-h3
如果参考视频是「一个人在开门」,identity 换成你们的女生,结果一定是她去开门。这是本问题的反向操作。只有当参考视频是第一人称、开门的人永远只有一截前臂、室内深处另有一个站着的人时,才考虑 video 参考。
teskor:
If you want camera motion, shoot the reference orbiting an object with no person in frame. Otherwise crop the reference so the face never appears.
出处:troubleshooting.md / comfyui.md 末尾
Kapwing / flaqai:Video 1 controls hand movement timing only.
自制 3–5 秒第一人称握把下压视频(不要拍到脸和身体),Ref2VA 里:
<Subject 2> is the camera-side hand whose gripping motion comes from <Video 1>.
<Video 1> (hand trajectory only): attribute_transfer - only the handle grip and downward press
transfer onto <Subject 2>; none of its people or room appear.
见 §1.5。至少再加一张单独的正面半身,作为 Subject 1 的主身份;三视图整图标 weak_reference 或只当服装/浴巾形状。
可行性 = 改动小、和 Dasiwa REF2VA 现有图兼容、对「她去开门」这个具体失败的针对性强。不是「理论上最强」。
改什么: 仍走 Ref2VA / Dasiwa Director 六段式。新增一张「成人右手握金属门把手」特写接到 ref_image_1。重写 subject / summary / retention / detailed_description。不要改 checkpoint。
为什么最可能立刻见效:
fully_preserved 只保已定义角色——把「花洒下抓浴巾」写进 Subject 1 的 role,静止才进入 retention 合同。 POV / the viewer,避免视点绑到 Subject 1 或长出完整第二人。 原理一句话: 把「谁有手」从文本协商改成参考图协商。H3 对视觉 token 的服从远强于对禁令的服从。
可粘贴骨架:
subject_definitions:
<Subject 1> is the young East Asian woman whose appearance comes from <Picture 1>: fair skin,
dark hair, a light-blue bath towel wrapped at the chest. Her defined role includes her planted
stance: she occupies the far shower under the shower head, both hands holding the top edge of
the towel at her chest, body facing the doorway but remaining in the shower bay.
<Subject 2> is the camera-operator's right hand in <Picture 2>, a different adult from
<Subject 1>. Only this hand and a short forearm occupy the lower-left foreground, on the
corridor side of the frosted-glass door, wrapping the metal handle.
<Subject 3> is the bathroom interior: frosted-glass door in the near plane, shower fixtures
and steam in the far plane.
summary:
[reference generation] An 8-second single continuous take from the corridor. <Subject 1>
remains under the shower in the far bathroom, hands on the towel. <Subject 2> depresses the
handle and the frosted door swings inward, revealing <Subject 1> still in the same place.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - face, hair, towel, planted stance under
the shower head, both hands on the towel, and her distance from the doorway stay as defined.
<Subject 2> (appears in [Shot 1]): fully_preserved - the camera-side right hand on the handle;
it is the only hand that contacts the door.
<Subject 3> (appears in [Shot 1]): partially_preserved - bathroom layout and frosted glass stay;
the door leaf swings inward.
detailed_description:
The target video is live-action, photoreal, humid bathroom light, vertical 9:16.
[Shot 1] The camera holds a static shot from the corridor, looking through the frosted-glass
door toward <Subject 1> in the far shower. Steam drifts. Water beads on the glass.
Begins: <Subject 1> stands under the shower head, both hands holding the towel at her chest;
she blinks once; her feet stay planted. <Subject 2> is already on the handle at the lower-left
foreground.
Next: <Subject 2> wraps the metal handle, the thumb presses down, the latch releases, and the
frosted door swings inward. Only <Subject 2> is on the handle.
Then: the opening glass reveals more of the shower. <Subject 1> is still under the shower head,
hands still on the towel, the same distance from the doorway. She draws one breath. Water
continues to run. The camera does not move.
Finally: the door stays ajar. <Subject 1> holds the same stance. <Subject 2> rests on the
handle.
overall_soundscape:
Bathroom extractor hum and shower spray continue. A short metallic latch click, then the soft
sweep of frosted glass. Light fabric tension from the towel. One quiet breath.
non_diegetic_music:
N/A
配套参数: ref_image_size=max(定稿)、seed 固定、length 124 或 158、关掉 turbo/cache、denoise 1.0。三视图若必须保留,加一张单独正面半身当 Picture 1,三视图整图降为弱参考。
风险: 没有手部参考图时,Subject 2 仍可能被画成她的手。手入画且朝镜头时 teskor 警告多指,写 five clearly separated fingers, a firm grip,并让手停在画面下缘,不要伸到画面中心。
改什么: Dasiwa 切到 I2VA/FL2VA(fl2va 权重)。用图像工具或一次静帧生成做出两张构图图(不必是真实摄影):
为什么排第二:
fl2va。 原理一句话: 把「她始终在深处」从文本条件升级为每步重注入的像素条件。
提示词要点: 官方对齐句必须是第一行;正文只写门如何从首帧摆到尾帧;写 settles into the pose, spacing, and composition established by Picture 2。不要写 POV。
可叠加 MiniMaxH3AddGuide,把同一张「她在花洒下」的静帧锚到中间帧,进一步防中段走位。
风险: FL2VA 不吃角色三视图那套 <Subject> 标签,身份一致性可能略差;两张构图图的光线/颗粒必须接近,否则会变成「美颜 morph」(teskor templates.md)。门的运动要是两帧之间能插值的,不要首帧关门、尾帧人已经换房间。
改什么(三选一,强度递增):
为什么排第三: 对「身体不要去开门」几乎是硬约束,但 Dasiwa 默认 REF2VA 图要接新节点、新模型(minimax_h3_fun_controlnet_union_*.safetensors)或二采图,排查成本最高。
原理一句话: 骨架/掩膜是空间约束,不走「文本里谁是主语」这条会串的通道。
风险: pose 视频如果把手也画成去握门,问题会原样回来——pose 里她的手必须一直在胸口。拆镜会失去「一镜从门外推开看到她」的连贯感,需要剪辑补。
| 做法 | 原因 |
|---|---|
| 继续加长六段式、反复写「她不做开门」 | CFG=1,禁令加权重;你们已经 9 次 |
使用官方运镜词 POV / first-person POV | 官方定义为 subject's point of view |
写 the viewer opens the door | Reddit:容易长出完整第二人 |
video editing 把女生替换进「有人开门」的视频 | image=identity, video=performance,她会去演开门 |
指望 ref_strength / 更高 steps / 换 seed 解决归属 | 这些不改变默认施事者 |
| 只丢三视图、不给手部图 | grid 还在和场景抢内容,且开门的手没有视觉源 |
| 把开门动作在 summary、retention、detailed_description 各写一遍 | 词频让开门赢 |
官方社区(minimaxh3.studio / teskor)都要求每次只改一个控制变量。
ref_image_size=max。 这是方案 1 的完整形态。 h3_prompt_guide.txt(身边有物=交互;首尾帧锁空间)Ref2VA 不会「理解」第一人称推门 vs 室内站着的人;它会把参考图里那个有手的人,派去完成 prompt 里最具体的那组手部动作。要分开这两件事,必须让开门的手拥有自己的 Subject 和自己的像素,或者用首尾帧 / pose / 掩膜在像素层锁住女生的身体。继续在六段式里用中文或英文强调「她不能开门」,官方 retention 语义和 CFG=1 都不会帮你执行这条禁令。