结论先行:这句话说的不是一件事,而是四件可独立验证的事——① 换一个针对"中景模糊/油腻皮肤/运动模糊"做过微调的 H3 模型(Singularity);② 用官方默认的 0.4MP 小画布(16:9 = 864×480)跑第一遍;③ 在两遍采样之间把视频 latent 放大 1.5 倍(864×480 ×1.5 = 1296×720,按 32 像素网格对齐后落到 1280×704,即 0.90MP);④ 用 VDN(线性注意力)和 NVIDIA Sol-Attn(稀疏注意力)做加速。
其中 ①②③ 是工作流层的做法,④ 是换内核层的加速,而"解决中景模糊"是模型层的固有属性(Singularity 的 HDR 微调),不是任何工作流参数能修出来的。
关键现实核对(针对你的 GPU 机 192.168.31.31,RTX 5060 Ti 16GB):
- ①②③ 的零件几乎全都在机器上了:Director 默认就是 0.4MP/16:9/864×480;学习型 H3 潜空间放大器(LBH 3D fp16/bf16)权重已下载;
ComfyUI-MinimaxH3AudioT8有MiniMaxH3LearnedLatentUpscaleT8Advanced(两遍二采专用),user/default/workflows/里已经躺着 6 个两遍采样工作流,其中一个的 target 就写着 1280×704。- 真正缺的只有两样:Singularity 权重(21 GB,hf-mirror 可下)和 VDN-H3 那套(节点 + int8 stage ≈ 2.6 GB)。
- "VDN + SOL 同时提速"有硬冲突:VDN 节点会接管
blocks.*.attn.forward,与"MiniMax H3 Scheduled Sol Attention"补丁不能叠加(叠加后跑的不再是 VDN-H3)。机器上现在装的是ComfyUI-SolAttn_triton(社区 Triton 版 Sol-Attn,就是"NVIDIA-SOL"),所以要么二者选一,要么换 xmarre/Sol-H3 + VDN-H3-Plus(provider API v3)那一套才能"组合"。
| 话里的词 | 到底指什么 | 依据 / 现状 |
|---|---|---|
| Singularity | Minimax-h3_Singularity,WarmBloodAban(AIGC-Singularity 团队)发布的 MiniMax H3 社区融合微调:把 ref/fl/b25-49 等多个 H3 变体合并成一个 ComfyUI 可直接加载的文件,再做三天深度微调+剪枝去掉高步数训练的伪影 | ComfyUI Wiki 新闻、HF 仓库。文件:Minimax-h3_Singularity_ref2va_Pruned_v1.3_int8.safetensors 20.97 GB(推荐)/ 完整版 34.0 GB |
| "很好地解决了中景模糊" | 官方列出的改动里明确写着:HDR 训练以减少高速动作的运动模糊、修复中景与远景镜头中折叠/模糊的远景人脸、去掉厚重"过度油光"皮肤、强化剑斗/武术动态、VFX 调优、运镜响应更好 | 同上表格式对比。这一条是模型属性——挂上这个 UNET 就生效,不需要任何节点 |
| 0.4 分辨率 | 官方 MiniMax H3 模板的默认画布 0.4 MP,16:9 下 multiple=32 取整 = 864×480。你自己 Director 插件的 README 原文:画布默认 0.4MP 16:9(864×480);minimax_h3_director_external_groups_r2v.json 里写死的 "megapixels":0.4,"width":864,"height":480 | 已在机器上核实(ComfyUI_MiniMaxH3_Director/example_workflows/README.md、web/js/minimax_gen_timeline.js) |
| 潜空间 1.5 | 第一遍去噪完成后,把 H3 的视频 latent 空间放大 1.5 倍,再跑第二遍。H3 视频 latent 空间压缩 = 16 像素/格且 DiT patch 是 (1,2,2),所以 latent 宽高必须为偶数、像素宽高必须能被 32 整除 | 参考 rockerBOO/h3-latent-upscaler(纯插值,snap 到偶数)与 LBH-123-AI/Minimax_h3_latent_Upscaler(学习型 3D latent 放大网络,691 MB,机器上已有 fp16/bf16 两份) |
| 双采样 | 就是图像圈的 hi-res fix / two-pass:第一遍低清把构图、运动、音频定下来 → 放大 latent → 同步放大 conditioning 里的参考图/首尾帧 → 第二遍低去噪(denoise 约 0.2–0.4)重采样补高频 → 音频单独按自己的 sigma 曲线加噪后回拼 → 解码 | seedance 两阶段指南、rockerBOO README、机器上的 2026-08-21_H3_Learned_Latent_TwoPass_*.json |
| 1280/704 | 0.4MP 画布放大 1.5 倍后的落地尺寸:864×480 = 0.41MP,×1.5² = 0.90MP ⇒ 1280×704(0.90MP 精确值)。算术见下节 | 机器上的 2026-08-21_H3_Learned_Latent_TwoPass_I2VA_Standard_Advanced_EXP.json 里,MiniMaxH3LearnedLatentUpscaleT8Advanced 的参数正是 … 'scale_by', 2.0, 1.0, 1280, 704, …——target 1280×704 是这台机器已经跑过的尺寸 |
| VDN | Video Delta Net(OpenVDN):MiniMax-H3 的混合注意力——近邻帧保留精确 softmax 窗口,远距离时序上下文走 checkpoint 里的 Video Delta Attention 线性分支(常数代价递归状态),注意力开销随片长线性而非平方增长;发布含 8 步 DMD 与 50 步两档 | OpenVDN/vdn-minimax-h3、ComfyUI 移植 Saganaki22/ComfyUI-VDN-H3;机器上未装 |
| NVDIA-SOL(NVIDIA Sol-Attn) | NVIDIA NVlabs Sana / Sol-Attn:免训练稀疏注意力(tau 阈值 + 池化质心路由),论文 arXiv 2607.24027 | 机器上已装:custom_nodes/ComfyUI-SolAttn_triton(社区 Triton 实现,README 自述已在 RTX 4090/5090 上对 MiniMax H3 测过)。B 站"Sol-H3 开源提速 15.5 倍"指的是理想配置下的数字 |
| "提速"的可信度 | ⚠️ VDN 官方"74.5×"= 8×B200 + FA4 + FP8 + 8 步蒸馏的总和;官方自己的单卡数字是 50 步下约 2.6×;这个 ComfyUI 移植版在 RTX 5090 上 1280×736/145 帧/8 步实测约 17 s/it,且作者自己写明"显存/内存有限就别用这个"(额外占 4.3 GB bf16 / 2.2 GB int8 分支权重) | VDN-H3 README 的 Hardware reality check 一节 |
| 阶段 | 像素 | latent(÷16) | 说明 |
|---|---|---|---|
| 第一遍画布(0.4MP, 16:9) | 864 × 480 = 0.41 MP | 54 × 30 | 32 的倍数 ✓、latent 偶数 ✓,官方默认 |
| latent ×1.5 | 1296 × 720 | 81 × 45 | 1.5 倍后 latent 变奇数,必须对齐 |
| 对齐到 32 像素网格 | 1280 × 704 = 0.90 MP | 80 × 44 | 1296/32 = 40.5 → 40、720/32 = 22.5 → 22(四舍六入五取偶正好落这里);用"向上取到偶数"的实现则是 82×46 = 1312×736 |
| 第二遍画布 | 1280 × 704 | 80 × 44 | 面积 = 0.9 MP = 0.4MP × 2.25 = 1.5² ✓ 完全自洽 |
要在工作流里把尺寸钉死成 1280×704,有三种写法(任选其一):
target_dimensions(T8 Advanced 节点有 target_width/target_height,step=32),直接填 1280 / 704;scale_by ≈ 1.4667(54×1.4667 = 79.2→80;30×1.4667 = 44.0 恰好偶)→ 精确 1280×704;MiniMaxH3LatentUpscaleBy32T8 的 best_aspect 策略)让它自己挑最接近比例的 32 倍数尺寸。出处备注:机器上
MiniMaxH3LatentUpscaleBy32T8的默认就是['bicubic', 1.5, '16 - MiniMax H3', 'best_aspect']——"潜空间 1.5"这个说法很可能就是从这条默认参数来的。
| 类别 | 具体 |
|---|---|
| 底座 | ComfyUI 0.34.0(commit 2026-09-02)、torch 2.11.0+cu130、triton 3.6.0、RTX 5060 Ti 16 GB(compute_cap 12.0 = SM120,Blackwell ✓) |
| H3 主模型 | minimax_h3_ref2va_pruned_int8_convrot 21.0 GB、minimax_h3_fl2va_pruned_int8_convrot 21.0 GB、*_pruned_w4a8_mixed 11.8/12.5 GB、dasiwa_minimax_h3_ref2va_..._8turbo_int8 21.0 GB |
| 文本编码器 | qwen3vl_32b_minimax_h3_nvfp4_awq、int4_convrot、heretic_nvfp4 |
| VAE | minimax_h3_video_vae_int8_convrot / fp16、minimax_h3_audio_vae_fp32 |
| Turbo LoRA | minimax_h3_ref2v_turbo_4step_v0.1、minimax_h3_fl2v_turbo_4step_v1.0_768p、fl2v_turbo_8step_v1.0、h3/minimax_h3_turbo_v4_step600_ema_pruned |
| 加速内核 | ComfyUI-SolAttn_triton(NVIDIA Sol-Attn)、ComfyUI-Spectrum-MiniMax-H3、sageattention 2.2.0(按 2026-09-06 运维结论已全局停用)、ComfyUI-INT8-Fast |
| 潜空间放大 | models/latent_upscale_models/minimax_h3_latent_upscaler_3d_fp16.safetensors + _bf16(学习型放大器权重已就位);ComfyUI-MinimaxH3AudioT8 提供 MiniMaxH3LearnedLatentUpscaleT8Advanced、MiniMaxH3LatentUpscaleBy32T8、MiniMaxH3TwoPassDetailMixerT8Advanced、MiniMaxH3DualClockSamplerT8 |
| 二采编排 | ComfyUI_MiniMaxH3_Director 的 MiniMaxH3DirectorRefine(upscale / h3_latent / 权重 / 采样器 / cfg / size_mode / 画幅 / scale_by / target_w / target_h) |
| 现成工作流 | 2026-08-21_H3_Learned_Latent_TwoPass_{I2VA_Standard,I2VA_Native_Speech,Hybrid_Lock_Source,Hybrid_Reference_Only,Hybrid_Remix_Source_020}_Advanced_EXP.json(target 1280×704)、2026-08-16_H3_Latent_Upscale_By32.json、minimax_h3_director_二采_加速.json、MiniMax_H3_I2V_10s_0.7MP_加速放大_修正版.json |
| 缺口 | 体积 | 下载源(hf-mirror 可达,实测 HTTP 200) |
|---|---|---|
| Singularity 模型 | 20.97 GB | WarmBloodAban/Minimax-h3_Singularity → Minimax-h3_Singularity_ref2va_Pruned_v1.3_int8.safetensors |
| VDN-H3 节点 | 仓库很小 | git clone https://github.com/Saganaki22/ComfyUI-VDN-H3(github 实测 HTTP 200) |
| VDN stage 权重(int8) | 2.30 GB 分支 + 1.18 GB 适配器 | drbaph/vdn-minimax-h3-int8-convrot-comfyui(bf16 原版分支 4.28 GB 也可,但 16GB 卡建议 int8) |
| (可不装)rockerBOO 插值放大节点 | — | 功能与 MiniMaxH3LatentUpscaleBy32T8 / T8 Advanced 重叠,没必要再装 |
磁盘:GPU 机 / 还剩 291 GB,放得下。
这是最短路径,只下载 Singularity 一个文件就能得到"这句话"里 95% 的效果(Singularity + 0.4MP + 潜空间 1.5 + 双采样 + 1280×704 + Sol-Attn 提速),唯一没有的是 VDN。
步骤
HF_ENDPOINT=https://hf-mirror.com ~/ComfyUI/venv/bin/hf download \
WarmBloodAban/Minimax-h3_Singularity \
Minimax-h3_Singularity_ref2va_Pruned_v1.3_int8.safetensors \
--local-dir ~/ComfyUI/models/diffusion_models
# 没有 hf CLI 时用 curl -L -C - 直接拉:
# https://hf-mirror.com/WarmBloodAban/Minimax-h3_Singularity/resolve/main/Minimax-h3_Singularity_ref2va_Pruned_v1.3_int8.safetensors
⚠️ 文件名带 ref2va ⇒ 参考生视频(r2v)任务;t2v/i2v 请用 fl2va 系底座,两者别混接。
minimax_h3_director_二采_加速.json 为模板
UNETLoader 换成 Minimax-h3_Singularity_ref2va_Pruned_v1.3_int8.safetensors;CLIPLoader type = minimax,用 qwen3vl_32b_minimax_h3_nvfp4_awq(或 int4_convrot 省显存);aspectRatio = 16:9、megapixels = 0.4、multiple = 32 ⇒ 面板显示 864×480;LoraLoaderModelOnly 挂 minimax_h3_ref2v_turbo_4step_v0.1(Singularity 官方就是推荐配这个 4-step LoRA 快速迭代);第一遍 4 步 / cfg 1.0 / shift video 12, audio 3;MiniMaxH3DirectorRefine(二采):mode = upscale、method = h3_latent、权重 minimax_h3_latent_upscaler_3d_fp16.safetensors、scale_by = 1.5(或 target 1280 / 704)、采样器 euler、第二遍 denoise 0.25–0.40、3 步、音频 denoise = 0(音轨沿用第一遍原声);PathchSageAttentionKJ 保持 auto(你的运维结论是 Sage 全局停用,这行按现状保留/绕过)+ 在 UNET 上加 Sol-Attn 节点(start_percent 0~0.1、end_percent 0.9~1.0、tau 从默认起调,先和关掉做 A/B)。2026-08-21_H3_Learned_Latent_TwoPass_I2VA_Standard_Advanced_EXP.json,把它的底座换成 Singularity、把 MiniMaxH3LearnedLatentUpscaleT8Advanced 的第一个 widget 从 scale_by=2.0 改成 1.5(target 保持 1280×704),MiniMaxH3TwoPassDetailMixerT8Advanced 的二采强度按抗糊档从 0.25–0.4 起调。参数起点表
| 参数 | 第一遍 | 第二遍 |
|---|---|---|
| 画布 | 864×480(0.4MP) | 1280×704(0.90MP) |
| 步数 / 采样器 | 4(turbo LoRA)/ res_multistep 或 euler | 3 / euler |
| denoise | 1.0 | 0.25–0.40(>0.5 会改动作) |
| sigma shift | video 12 / audio 3 | video 12 / audio 3(音频分支走 MiniMaxH3ShiftSigmas 12→3) |
| 音频 | 生成 | denoise = 0,沿用第一遍 |
| 放大 | — | 学习型 3D latent 放大 ×1.5 |
cd ~/ComfyUI/custom_nodes && git clone https://github.com/Saganaki22/ComfyUI-VDN-H3
HF_ENDPOINT=https://hf-mirror.com ~/ComfyUI/venv/bin/hf download \
drbaph/vdn-minimax-h3-int8-convrot-comfyui \
--local-dir ~/ComfyUI/models/vdn/vdn-minimax-h3-int8-convrot-comfyui
节点用法:插在 H3 UNETLoader 与采样器之间,apply_turbo_adapter=ON ⇒ 采样步数必须是 8(对应 stage-dmd),lora_mode=merge(8 步 DMD 档必须 merge,bypass 会出颗粒/劣化),branch_weights=auto。
ComfyUI-SolAttn_triton 的 Sol 补丁(两者都改写 blocks.*.attn.forward,叠加后 VDN 的线性分支被跳过,而且 VDN 的 LoRA 会被用在它没训过的注意力上);cute_sm120 后端)+ VDN-H3-Plus 的 provider API v3,才能"Sana SOL 处理 VDN 的局部 Q 行、K/V 域保持 VDN 语义"地组合。代价:需要 triton>=3.6,<4(你已满足)、NVIDIA CUTLASS DSL(CUDA 13 extra)、cuda-python、apache-tvm-ffi,且仅限 Linux + SM120(你这台正好符合,但会改动现有环境)。NestedTensor AV latent,用原生 LTXVSeparateAVLatent / LTXVConcatAVLatent 拆合(已确认在你 0.34.0 的 comfy_extras/nodes_lt.py 里)。AddNoise:① 它内部对 NestedTensor 调 count_nonzero 会直接崩;② 即使喂普通 tensor,对 H3 这种 CONST 参数化流模型的输出也是错的(SamplerCustomAdvanced+DisableNoise 会再乘一次 (1-sigma),实测 denoise=0.4 解出来是纯噪声/棋盘格)。必须用 MiniMaxH3AddNoise(做了 inverse_noise_scaling 预抵消)。shift_video=12 / shift_audio=3 换算后期望音频噪声是 ≈0.14。所以视频分支用原始 sigmas_pass2,音频分支要串一个 MiniMaxH3ShiftSigmas(12→3)再进 MiniMaxH3AddNoise。MiniMaxH3ConditioningUpscale(或 T8 的等价节点),scale_by/upscale_method 与 latent 放大保持一致。VAEDecode 一次性解全部帧会 OOM ⇒ 用 tiled VAE decode(保持默认 overlap,有缝就调大 tile 而不是 overlap)。17k+5 网格(124 / 141 / 227…),画布宽高是 32 的倍数,latent 宽高是偶数——违反就报 shape/patch 错。| 步 | 动作 | 预估 |
|---|---|---|
| 1 | hf-mirror 拉 Singularity Pruned int8(20.97 GB)→ models/diffusion_models/;校验 sha256 / 文件大小 | 35–70 min(后台跑) |
| 2 | 复制 minimax_h3_director_二采_加速.json 为 ..._Singularity_0.4MP_1.5x.json,换 UNET、画布 0.4MP/16:9、二采 h3_latent ×1.5(target 1280×704)、挂 Sol-Attn | 15–30 min |
| 3 | 首跑 5s/124 帧 基准;A/B:Sol-Attn on/off、二采 on/off | 每条约 10–25 min(取决于是否 offload) |
| 4 | 达标后再拉长到 10–15s;若要长片再评估 VDN(注意与 Sol-Attn 二选一) | — |
| 5 | 产出/结论写回 dl-hub + 运维记录 | — |
ComfyUI_MiniMaxH3_Director/example_workflows/README.md(0.4MP 默认画布)、user/default/workflows/2026-08-21_H3_Learned_Latent_TwoPass_*.json(target 1280×704)、dl-hub/06-ComfyUI-H3运维记录/2026-09-06-H3提速方案.md本报告写完后已在 GPU 机把路线 A 跑通,两条关键结论回填:
ComfyUI_MiniMaxH3_Director/director/refine_pack.py 的 infer_upscale_target(864,480) → nw = round(864×720/480) = 1296、nh = 720 → snap_dimension() 按 32 取整(Python 四舍六入五取偶)→ 1296→1280、720→704。所以只要二采的"宽高比"设成 跟随导演台、一采是 0.4 MP/16:9(864×480),目标画布自动就是 1280×704。踩坑点:若改成 16:9 (宽屏) + megapixels,0.88 MP 得到 1280×736、1.5 MP 得到 1664×928,都不是 1280×704。落地产物见同目录 Singularity双采样工作流-交接说明.md(工作流改动、参数表、烟测与 A/B 数据、对比视频)。
报告生成:DSH 会话(GPU 机 192.168.31.31 实测盘点 + 公开资料核实)。"已有/缺失"栏均为远程 ls/grep/pip list 实测结果。