c194dd00(2026-09-21 04:57 UTC 拉取)| 问题 | 结论 | 依据 |
|---|---|---|
| 是 ComfyUI 本体适配,还是节点更新? | "能让 H3 显存/速度跳一档"的能力全部来自 ComfyUI 本体(内核)适配,而且都是 2026-09-05 ~ 09-20 合并的,不是这两天的自定义节点更新 | 见第二节提交时间线 |
| 「今天」有新适配吗? | 没有。 09-21 ~ 09-24 上游没有任何 H3 显存/速度相关提交;这三天 H3 只有两个补丁:VAE 瓦片拼接、rms_rope 崩溃修复 | 见第二节表 |
| 「掉了 10G 显存」能由注意力机制解释吗? | 不能。 本机实测 comfy-kitchen INT8 注意力显存不降反升(8K token:稠密 86 MiB → INT8 617 MiB)。10GB 级别的下降只可能来自 权重换成更小的量化档 或 ComfyUI 缓存被释放 | 见第三、四节实测 |
| 「速度瞬间提很多」能解释吗? | 能,且量级吻合。 H3 5s@1MP ≈ 2~3 万 token,正处于本机实测的交叉点之后:32K token 时稠密 479.8 ms → INT8 244.6 ms(1.96×);社区 SPEED V2 渐进分辨率节点自报 1.23–2.41× | 见第四节 |
一句话:如果网友没动工作流却"变快了",那是内核适配(Comfy 编译器 / 核心稀疏注意力 / comfy-kitchen INT8);如果显存真的少了 10G,那几乎只能是换了更小的量化权重(本机两档 H3 权重相差 8.43 GB)或者缓存被释放(本机实测空闲常驻 13.78 GB,一条 /free 就掉到 310 MiB)——与"注意力节点更新"无关。
2026-09-24 03:01 1568e6cf Lower memory usage and .comfy_attention support for lumina family models. (#16515)
2026-09-24 01:57 3b4c0b0e feat: ming-image support (#16482)
2026-09-23 03:16 b5cc8830 Port some optimizations to flux model family. (#16488)
2026-09-23 01:39 912fca4f Fix MiniMax-H3 VAE rms_rope crash on offloaded qk_norm_scale (#16485) ← H3,仅修崩溃
2026-09-22 16:27 95539f56 support union cn 2.0 (#16471)
2026-09-22 06:22 b33e2b55 workflow templates v0.11.68 (#16466)
2026-09-22 01:57 e638023d Fast disk detection to all model loaders. (#16425)
2026-09-22 01:56 fc584aaa Blend H3 VAE tiles against composited neighbours (#16436) ← H3,仅 VAE 拼接画质
2026-09-22 01:49 d1584209 Cuda graphs + memory compiler on ace step 1.5 (#16461)
2026-09-21 05:58 b0f4b7b2 JsonExtractString ...
2026-09-20 17:31 c194dd00 Allow model files to contain which attention should be used for each block. (#16419) ← 关键(本机 HEAD)
2026-09-20 15:27 73c9bad4 ComfyUI v0.37.0
2026-09-20 02:50 c8ed2c8c Lower wan peak vram when using comfy kitchen attention. (#16418)
结论:09-21 之后没有一条 H3 显存/速度提交。H3 最近真正碰显存的是 09-15(f14bbe28 Lower minimax VAE usage by a bit、b2e31e89 MiniMax-H3 VAE optimizations)。
真正能造成"跳一档"的内核改动,全都早于"今天":
| 日期 | 提交 | 内容 | 作用侧 |
|---|---|---|---|
| 09-05 | #15861 | Comfy 编译器(comfy-aimdo 内存编译器 + CUDA graphs),Wiki 明说让 H3 720p/158 帧从"停滞"变为 ~43 s/step,并让"物理显存峰值=逻辑分配峰值,不再让已释放分配留在缓存" | 本体 |
| 09-06 | #16072 | 核心「Block Sparse Attention」节点(Sol-Attn / top-k SLA / VSA 三后端,H3 首个转换) | 本体(核心节点) |
| 09-07 | #16148 | 长时稀疏分配时暂停 Comfy 编译器(修早期 OOM/崩溃) | 本体 |
| 09-08 | #16154 | 统一模型注意力节点(Model Attention Backend,可选 comfy kitchen attention = INT8) | 本体(核心节点) |
| 09-20 | #16419 | 允许模型文件逐块声明用哪种注意力(本机 HEAD) | 本体 |
attention.config 张量(本体适配,本机未生效)本机 comfy/ldm/modules/attention.py:76 的 ComfyAttention:
metadata = state_dict.pop(prefix + "config", None) # 从权重里取 "….comfy_attention.config"
if metadata is not None:
config = json.loads(metadata.numpy().tobytes())
method = config.get("attention") # 支持 "comfy_kitchen_int8"
if method == "comfy_kitchen_int8" and comfy_kitchen.int8_attention_is_available(...):
self.function = attention_comfy_kitchen_int8 # 自动切 INT8 注意力
Attention.forward() 已经把 preferred_attention=self.comfy_attention 传下去(comfy/ldm/minimax/model.py:161,200)——这是 #16419(09-20)本体适配。MiniMaxH3/ 子目录 3 个)。见第四节实测:INT8 显存高于稠密。所以"掉 10G"不能归因于注意力后端。
本机现成两档 H3 权重实测文件大小:
| 权重 | 大小 | 差值 |
|---|---|---|
minimax_h3_ref2va_pruned_int8_convrot.safetensors | 20.97 GB | — |
minimax_h3_ref2va_pruned_w4a8_mixed.safetensors | 12.54 GB | −8.43 GB |
w4a8_int8_linear、int8_linear、quantize_int8_convrot_weight、sol_attn 等),不需要第三方节点。本机实测(空闲、队列为空):
释放前:nvidia-smi memory.used = 13782 MiB (进程 5056 占 13756 MiB)
POST /free {"unload_models":true,"free_memory":true} → http 200
释放后:nvidia-smi memory.used = 310 MiB
Model QwenImage21TEModel_ prepared for dynamic VRAM loading. 8916MB Staged.)。--cache-lru、--fast-disk、动态 VRAM 策略)会让这个数字突然掉十几 G,但生成速度未必变化;反过来 Comfy 编译器/aimdo(09-05)宣称的正是"不再让已释放的分配留在缓存里"→ 视觉上就是"显存占用掉了",同时步进更快。StanLukuvka/ComfyUI-MiniMax-H3-SPEED(custom_nodes),把早期去噪步骤放在低分辨率网格上跑,再逐步放大。H3 注意力维度从权重头推得:blocks.0.attn.qkv_proj.weight [21504, 2688]、out_proj [5376, 3584] ⇒ hidden 3584 / 56 heads / head_dim 96。 GPU:RTX 5060 Ti,sm_120;comfy-kitchen 0.2.35;int8_attention_is_available=True、sol_attn_is_available=True。
| 序列长度 S | 稠密 PyTorch SDPA 显存 | INT8 显存 | 稠密耗时 | INT8 耗时 | 加速 |
|---|---|---|---|---|---|
| 8,192 | 86 MiB | 617 MiB | 33.4 ms | 42.3 ms | 0.79× |
| 16,384 | 172 MiB | 1,233 MiB | 119.8 ms | 62.0 ms | 1.93× |
| 32,768 | 343 MiB | 2,466 MiB | 479.8 ms | 244.6 ms | 1.96× |
| 65,536 | 未测 | 4,932 MiB | — | 942.7 ms | — |
| 98,304 | 未测 | 7,398 MiB | — | 1,986.0 ms | — |
要点:
frames_per_token=(1,4,4,4,4)。5s、约 1 MP(如 768×1344 / 1280×720)落在 2~3 万 token 量级 ⇒ 正好是上表 INT8 已有 2× 收益的区间,与"速度瞬间提很多"量级相符。| 项 | 值 |
|---|---|
| ComfyUI 版本 | v0.37.0,HEAD c194dd00,拉取于 2026-09-21 04:57 UTC(git reflog) |
| 上游最新(核查时) | 09-24 1568e6cf;tags 已有 v0.37.1 / v0.37.2 |
| comfy-kitchen | 0.2.35(cuda 后端可用,含 sol_attn、w4a8_int8_linear、int8_linear) |
| comfy-aimdo | 0.5.5(日志:comfy-aimdo inited for GPU: NVIDIA GeForce RTX 5060 Ti) |
| 启动参数 | --lowvram --force-fp16 --reserve-vram 1.0 --fast-disk --cache-lru 20 |
| 稀疏/注意力文件 | comfy_extras/nodes_sparse_attention.py 存在;ModelAttentionBackend 节点存在 |
| 官方 H3 Director 工作流 | 12 个 JSON,均不含注意力补丁节点(仅 MotionContext-fl2va-ref2va.json 引用未安装的 H3SLAAttention) |
带 attention.config 的模型 | 0 / 84 |
cd ComfyUI && git log -1 --date=iso → 看拉取日期。若在 09-20 之前,就不可能有逐块注意力(#16419)。...int8_convrot(≈21 GB)还是 ...w4a8_mixed(≈12.5 GB)。显存"掉 10G"基本就是这一项。→ 用 ls -l models/diffusion_models/。Model Attention Backend 或 Block Sparse Attention 节点(本体核心节点,选了 comfy kitchen attention 就会提速 ~2×)。ComfyUI-MiniMax-H3-SPEED(渐进分辨率,节点侧提速)、ComfyUI-SolAttn*、ComfyUI-KJNodes 等。→ ls custom_nodes/。POST /free 试一次——本机实测能一次掉 13.5 GB,属于缓存而非真实占用。~/Downloads/v2ray-setup/refresh-sub.sh 刷新订阅 → 节点更新为 176.122.134.43 / 176.122.191.251 → 验证 Google 200、下载 4139 KB/s。POST /free。这只清缓存,不影响运行;下次出图会自动重新加载模型(首次略慢)。# 上游提交(走 v2ray 代理)
curl -s -x http://127.0.0.1:10809 \
"https://api.github.com/repos/Comfy-Org/ComfyUI/commits?since=2026-09-20T00:00:00Z&per_page=100"
# H3 注意力维度(从 safetensors 头)
# blocks.0.attn.qkv_proj.weight [21504,2688] / out_proj [5376,3584] -> 56 heads x 96
# 全库扫描 attention.config
python3 -c "..." # models/**/*.safetensors 84 个,命中 0
# 显存 A/B(GPU 机 venv)
cd /home/zyw/ComfyUI && venv/bin/python /tmp/h3_attn_bench3.py
# 空闲缓存实测
curl -s --noproxy '*' -X POST http://127.0.0.1:8189/free \
-H 'Content-Type: application/json' -d '{"unload_models": true, "free_memory": true}'