本文对应报告:
H3-Singularity双采样1280x704方案解析与落地.md执行时间:2026-09-11(DSH 会话落地)
| 项 | 内容 |
|---|---|
| 新工作流 | ComfyUI/user/default/workflows/minimax_h3_director_二采_Singularity_0.4MP_1.5x_1280x704.json |
| 来源模板 | minimax_h3_director_二采_加速.json(原文件未改动,可随时回滚) |
| 新模型 | models/diffusion_models/Minimax-h3_Singularity_ref2va_Pruned_v1.3_int8.safetensors(20.97 GB,aria2c 16 线程 + sha256 校验) |
| 模型来源 | hf-mirror:WarmBloodAban/Minimax-h3_Singularity(HF 官方域名在本网不可达,走镜像) |
| 已存在、直接复用 | H3 学习型潜空间放大器 minimax_h3_latent_upscaler_3d_fp16/bf16(models/latent_upscale_models/)、ComfyUI-SolAttn_triton(NVIDIA Sol-Attn)、Turbo LoRA、VAE、文本编码器 |
没有装:VDN-H3(路线 B;因为 VDN 与本机 Sol-Attn 会争抢 blocks.*.attn.forward,二选一,本次按"提速性价比优先"选 Sol-Attn)。
| 项 | 值 |
|---|---|
| UNET | Minimax-h3_Singularity_ref2va_Pruned_v1.3_int8.safetensors(ref2va ⇒ r2v / 参考生视频任务) |
| LoRA | minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensors(strength 1.0) |
| 画布 | Director 输出面板:16:9 / 0.4 MP / multiple 32 → 864×480 |
| 采样 | steps 8、res_multistep / simple、cfg 1.0、shift video 12 / audio 3(沿用模板;配 4 步 Turbo LoRA 时可降到 4 步,见"下一步") |
| 加速 | PathchSageAttentionKJ(auto) → MiniMaxH3MemoryEfficientSageAttentionPatch → SolAttnPatch #39 |
| 项 | 值 |
|---|---|
| mode / method | upscale / h3_latent(学习型 3D latent 放大,不是插值) |
| 权重 | minimax_h3_latent_upscaler_3d_fp16.safetensors |
| 目标画布 | 宽高比 = 跟随导演台 → 864×480 自动推 720P 档 → 1280×704 |
| 采样 | euler、passes 1、seed inherit;配对 BasicScheduler:beta / 3 步 / denoise 0.2 |
| 加速 | 第二条链:Turbo LoRA → Sage → SolAttnPatch #40 |
| 音频 | 沿用一采原声(Refine 内部按 shift 12/3 处理音频分支) |
Director 插件 director/refine_pack.py:
def infer_upscale_target(base_w, base_h): # 880行附近
if w >= h:
nh = 720
nw = max(32, round(w * nh / h)) # round(864*720/480) = 1296
return ensure_minimax_canvas(nw, nh) # snap_dimension: round(v/32)*32
# snap_dimension(1296) = round(40.5)*32 = 40*32 = 1280 ← Python 四舍六入五取偶
# snap_dimension(720) = round(22.5)*32 = 22*32 = 704
即 864×480 → 1296×720 →(×32 对齐)→ 1280×704。这句话里"0.4 分辨率 + 潜空间 1.5 → 1280/704"的数字,正是这条内置推导链算出来的(0.41 MP × 2.25 = 0.90 MP)。
注意:不要把宽高比设成
16:9 (宽屏)再填 megapixels——那条路走的是resolution_from_selector,0.88 MP 会得到 1280×736、1.5 MP 会得到 1664×928,都不是 1280×704。要钉死 1280×704 只有两条路:① 宽高比 = 跟随导演台(本次采用,自动推);② 宽高比 = 自定义 + 手填 1280 / 704。
在两条模型链的末端(Sage patch 之后、消费者之前)各插一个 SolAttnPatch:
一采链: UNETLoader#1 → LoraLoaderModelOnly#16 → PathchSageAttentionKJ#14
→ MiniMaxH3MemoryEfficientSageAttentionPatch#15 → SolAttnPatch#39 → MiniMaxH3Director#12
二采链: UNETLoader#23 → LoraLoaderModelOnly#32 → PathchSageAttentionKJ#30
→ MiniMaxH3MemoryEfficientSageAttentionPatch#31 → SolAttnPatch#40 → BasicScheduler#33 / DirectorRefine#38
Sol 默认参数(沿用机器上既有工作流的社区调参):tau=1.3, start_percent=0.2, end_percent=0.9, min_tokens=4096, int8_qk=True, sink_conditioning=exact_kv_and_rows, morton=False, int8_pv=True, use_tma=False, dense_blocks=""。
SolAttnPatch 两个节点右键 Bypass(Ctrl+B) 即可,无需改线。minimax_h3_director_二采_Singularity_0.4MP_1.5x_1280x704r2v_groups 需自己连一组),时长/提示词按需改h3_latent、权重 ...3d_fp16...BasicScheduler(二采)= beta / 3 / denoise 0.2常见坑(都已在本模板里避开):
skip_fl2v=True:首尾帧镜头默认跳过二采(保护钉死的锚帧)| # | 变量 | 起点 | 说明 |
|---|---|---|---|
| 1 | Sol-Attn on/off | 开 | Bypass 两个 Sol 节点,比对耗时与画质(Sol 论文定位:长序列/高分辨率收益更大) |
| 2 | 一采步数 | 8 → 4 | 已挂 4 步 Turbo LoRA,4 步是一采性价比档;看质量再决定 |
| 3 | 二采 denoise | 0.2 → 0.3 / 0.4 | 0.2 保守(身份稳),0.3–0.4 细节更多但可能改动作 |
| 4 | 放大幅度 | 1280×704 → 1280×736 / 0.94MP | 若想要更接近 16:9;改宽高比为 16:9 (宽屏) + megapixels 0.88 |
| 5 | 二采采样器 | euler → res_multistep | 插件作者提示高质量二采常用 res_multistep |
预期:16 GB 卡上一采 864×480 + 二采 1280×704 的峰值显存远大于单遍 864×480(二采像素是一采的 2.25 倍),首跑请留意 offload 抖动;显存不够时把底座换成 ..._w4a8_mixed(11.8 GB)或降二采目标。
minimax_h3_director_二采_加速.json 未被修改models/diffusion_models/Minimax-h3_Singularity_ref2va_Pruned_v1.3_int8.safetensors(删除可回收 21 GB)~/dl_singularity.sh,日志 ~/singularity_dl.log/tmp/sing_smoke_prompt.json + /tmp/sing_smoke_run.py(GPU 机):直接 POST http://127.0.0.1:8189/prompt,用 dasiwa V18 的 33 节点 API 展开图改造——
minimax_h3_ref2v_turbo_4step_v0.1BasicScheduler steps → 4ModelPatchTorchSettings#1512:2663 之后插 SolAttnPatch(verbose=true,便于看内核是否真的命中)→ 接 Director.ref2va_modelvideo/Singularity_smoke_0.4MP_4step_SOL用途:不依赖 UI,验证「Singularity 能加载 + 4 步 Turbo 能出片 + Sol 内核命中」。跑法:
nohup ~/ComfyUI/venv/bin/python /tmp/sing_smoke_run.py > ~/sing_smoke.log 2>&1 &
tail -f ~/sing_smoke.log
环境:RTX 5060 Ti 16 GB / torch 2.11.0+cu130 / Singularity Pruned int8(staged 19995 MB)+ ref2v Turbo 4-step LoRA 任务:dasiwa V18 REF2VA,864×480(0.4 MP)、294 帧、12 s、24 fps、cfg 1、shift 12/3、固定 seed 666666、4 步
Model MiniMaxH3 prepared for dynamic VRAM loading. 19995MB Staged. 208 patches attached[DaSiWa Advanced LoRA Loader] 'minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensors' mode=Basic full:624@1.00[sol_attn] chaining onto an existing attention override; Sol-Attn takes first refusal and delegates everything else to it
[sol_attn] conditioning sink: KV blocks (0, 67) exact, dense query blocks (0, 67)
[sol_attn] sparse (1, 39518, 56, 128) tau=1.3 int8 pointeroutput/video/Singularity_smoke_0.4MP_4step_SOL_00001_audio.mp4(294f 864×480@24fps,AV1/av1_nvenc + AAC)| 运行 | 采样耗时 | 端到端 | 说明 |
|---|---|---|---|
| Sol ON(冷启,含内核 autotune ≈25 s + 首轮预热) | 2:43(163 s) | 274.2 s | 每次重启 ComfyUI 后的首轮要付这个一次性开销 |
| Sol OFF(热) | 2:09(129 s) | 236.8 s | 对照基线 |
| Sol ON(热) | 1:52(112 s) | 218.9 s | 与基线同条件 → 采样提速 13.3%,端到端 7.6% |
e2a2500117e3d10521dcab71153afc64)→ Sol 路径确定性可复现;Sol OFF 输出不同(不同注意力数值路径),可肉眼 A/B。对比视频(下载中心):
DaSiWa_LTX2LoraLoader 现在多一个必填 use_cache,老模板直接报 Failed to validate prompt for output 2568 … Required input is missing: use_cache,而且它不会报错失败,只会静默"输出被忽略、0.5 秒就 success"——很容易误判成"跑通了"。curl -X POST http://127.0.0.1:8189/free -d '{"unload_models":true,"free_memory":true}' 后降到 332 MB 才够加载 20 GB 模型。本文件与报告同目录发布:http://192.168.31.76:8899/28-H3-Singularity双采样提速方案/