_clusterD.md

Cluster D:sm_120/Blackwell 兼容性坑 + 2026 新增实时数字人仓库

调研日期:2026-09-22。目标硬件:RTX 5060 Ti 16GB / sm_120(消费级 Blackwell)/ 32GB RAM / Ubuntu 24.04。 所有结论均附可点击来源。查不到的写「未查到」,不做推测。 抓取方式:curl 直取 README / GitHub HTML / shields.io 徽章;github.com 的 HTML 在调研后段一度不可达(返回 000),因此少数条目只保留了已验证的徽章值。


Section 1:sm_120 / Blackwell 兼容性坑(视频扩散 & 数字人栈)

1.1 总纲:消费级 Blackwell 没有 tcgen05 / WGMMA / TMEM

这是一切坑的根因。sm_120(GeForce RTX 50 系)用的是 Ampere 时代的 mma.sync.aligned.m16n8k16(MMAv2),没有 data-center Blackwell(sm_100/103)才有的 tcgen05、WGMMA、Tensor Memory。

实践含义(对 16GB 5060 Ti):任何「只在 sm_90a / sm_100a 上编译」的 attention 内核(FA3、FA4、SageAttention3 FP4)在 sm_120 上要么编译失败、要么静默走退化路径。


1.2 SageAttention on sm_120(坑最多,与本机 H3 掉总线问题直接相关)

(a) 官方 PyPI 包没有 sm_120 内核 → 静默回退 Triton JIT / 无加速

(b) 关键内幕:sm_120 的 dispatch 实际调用的仍是 SM89 内核(不是原生 sm120 内核)

(c) 具体报错清单(按严重度)

#现象 / 报错原文环境来源
1驱动级 GPU lost:No devices were found / GPU is lost(需重启);~54s 进入稳态生成时崩。原生 PyTorch attention 在更高显存(15.3GB vs 13.0GB)与更高功耗下跑通,排除显存/功耗/温度归因RTX 5060 Ti 16GB + MiniMax H3 + _sageattn_int8_fp8_nhd(int8-QK / per-channel FP8-PV)#391
2RuntimeError: CUDA error: misaligned address(sageattn3_blackwell 推理时)RTX 5060 Ti + PyTorch 2.12.0.dev+cu130#357
3ptxas error: Instruction 'wgmma.fence' not supported on .target 'sm_120';wgmma.mma_async with integer types / with FP8 types 同样不被支持(因 TORCH_CUDA_ARCH_LIST="9.0;12.0" 而误用 sm90 设置编 sm120)RTX 5090, CUDA 12.8, torch 2.8.0+cu128#291
4静默噪声:序列长度 >~160k token 时 FP8-PV 内核静默输出纯噪声(无异常、无 warning),采样器全步正常跑完;同配置关掉 SageAttention(SDPA)全部正常。~154k 以下干净,167k 起坏RTX PRO 6000 Blackwell(sm_120) + MiniMax H3#388
5CUDA Graph 重放静默错:capture 一次、之后往静态 buffer 写新值重放,输出相对 eager 偏差约 90%,无 NaN(下游解码成纯静态/噪声视频)RTX PRO 6000 Blackwell(sm_120)#392
6Error: Failed to initialize the TMA descriptor 1 → CUDA context illegal instruction(~73,774 token,per_block_mean=True;改 False 则正常)sageattn3==1.0.0 sm_120a, torch 2.13.0+cu130#382
7编译崩溃 NameError: name 'num' is not defined(setup.py 的 if/elif 链不覆盖 12.1)DGX Spark sm_121#330
8SageAttention3 FP4 在 sm_120 上几乎无收益:内核级仅 1.17×、端到端 0.98×(论文在 5090 上是 1.84×)RTX PRO 5000 Blackwell(sm_120)#378
9ComfyUI/KJNodes 侧报错原文:RuntimeError: sageattention is not new enough version or could not determine CUDA architecture, cannot apply MiniMax H3 Memory Efficient Sage Attention Patch.ComfyUI + MiniMax H3kijai/ComfyUI-KJNodes#721

本机自测记录的 sageattention/quant.py per_channel_fp8 OOM 这一具体报错文本,未查到对应的公开 issue 来源(GitHub HTML 调研后段不可达,无法穷举 issue 列表),故不列为有来源的结论;其根因与上表 #1/#4/#5 同源(sm_120 上走 SM89 内核路径)。

(d) 有 sm_120 内核的版本 / 已知修复

(e) 缓解建议(工程结论)

  1. 在 5060 Ti 上,长序列视频扩散(H3/Wan 等)优先关掉 SageAttention(用 PyTorch SDPA),#388/#391 都证明同一工作流 SDPA 正常。
  2. 如必须开:用 woct0rdho post6 / CUDA≥12.8 的 sm120 wheel,并把 pv_accum_dtype 调稳(sageattn_qk_int8_pv_fp16_cuda 最不易 overflow);见 woct0rdho README「Use notes」。
  3. 不要用 SageAttention3 FP4 换速度(sm_120 上无收益且更脆)。

1.3 FlashAttention on sm_120

版本sm_120 状态来源
FA2(flash-attn 2.x)预编译 wheel 不含 sm_120,直接报 RuntimeError: CUDA error: no kernel image is available for execution on the device。提问者反馈代码里最高只到 8.9;维护者答复编译器已指示支持 120(即需自行源码编译 + TORCH_CUDA_ARCH_LIST 含 12.0)Dao-AILab/flash-attention#1638
FA3依赖 Hopper WGMMA,在 sm_120 上不可用,需显式 guard 以防误走 Hopper 路径PR #2593「Guard Hopper FA3 path against SM120 devices」 · #1810「FA3 attention sinks blackwell sm120」
FA4(CuTe DSL)设计上走 tcgen05/TMEM → 原生不支持消费级 Blackwell;2026 年社区在补 sm120 路径(flash_fwd_sm120.py),但被 paged KV + num_splits 卡住,vLLM 服务化不可用;维护者原话倾向「不做」#2307 · 补丁:#2499 · #2634 · #2758 · #2771
FA4 sm120 实跑 bugflash_attn_varlen_func 在 ≥4 个 varlen 段、真实负载显存布局下 illegal memory access(GeForce Blackwell,测试于 MiniMax-H3 混合注意力视频模型,flash-attn-4 4.0.0b26/b29)#2860
其它 sm120/SM80 路径 bugpack_gqa 的 crd2idx 拒绝嵌套坐标;cu_seqlens[num_batch+1] OOB 读#2444 · PR #2862
Windows 编译CUDA 13.4 + MSVC 传统预处理器:fatal error C1189: MSVC/cl.exe with traditional preprocessor is used … pass /Zc:preprocessor;/std:c++17 与 torch 2.15 头文件(需 C++20)冲突;且 无 sm_121 gencode#2832

仓库基线:Dao-AILab/flash-attention(25k⭐,BSD-3-Clause,最后提交 2026-09 上旬)。

结论:在 5060 Ti 上不要指望 flash-attn;用 PyTorch SDPA 或(若必须)SageAttention2 的 sm120 wheel。


1.4 wmma / wgmma 指令不兼容报告汇总


1.5 PyTorch / CUDA 版本要求(sm_120)


1.6 xformers / Triton 版本约束


1.7 ONNX Runtime / CUDA EP 在 Blackwell 上的坑

⚠️ opset 相关:本次调研未查到「Blackwell 与特定 ONNX opset 版本冲突」的一手官方来源,故不列结论。


Section 2:2026 年(1–9 月)新增 / 新活跃的实时流式数字人仓库

判定口径:用 shields.io created-at 判「是否 2026 新建」,last-commit 判「是否有 2026 活动」;星数取 shields.io stars(2026-09-22 实测)。已排除题目给定的已知名单(除注明「2026 有重大更新」)。

(A) 2026 年全新仓库

仓库链接⭐License新建最后提交一句话实时/流式?
SentiAvatar/SentiAvatarhttps://github.com/SentiAvatar/SentiAvatar453GitHub 无法识别(论文 CC BY-NC-SA 4.0,见 arXiv 2604.02908)2026-042026-04SentiPulse(+ 人大 GSAI)开源的交互式 3D 数字人框架:SuSuInterActs 数据集(21K clip/37h) + 运动基础模型 + plan-then-infill + RVQVAE/Face VQVAE,输出 BVH / UE JSON(不是像素视频)✅ 实时:README 原文 "Generates 6 seconds of motion in 0.3 seconds with unlimited multi-turn streaming"
KlingAIResearch/AvatarForcinghttps://github.com/KlingAIResearch/AvatarForcing81GitHub 无法识别2026-032026-05快手 Kling 的一步流式 talking avatar:单参考图 + 音频 + 文本 → 视频;局部未来滑窗 + 异构噪声联合去噪,每步产出 1 个干净 block,恒定单步开销。学生模型 1.3B(Wan2.1-T2V-1.3B),论文实测 34 ms/frame、25 FPS、832×480✅ 流式(one-step streaming diffusion) — arXiv 2603.14331
PeterIverson/Super-Starhttps://github.com/PeterIverson/Super-Star10Apache-2.02026-072026-08[ACM MM 2026] "Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans"。3D 数字人在线实时交互:Streaming Speech Response(Qwen3-Omni-30B-A3B,需 CUDA driver ≥ 12.4)+ Online Gesture Generator(动作 RVQ-VAE + MotionGPT,torchrun --nproc_per_node=1 可单卡训练/推理);2026-08-14 放出训练/推理/评测代码,论文 arXiv 2608.24909,项目页 https://super-star-2026.github.io/✅ 流式实时交互(在线 pipeline,含 streaming speech + online gesture)
rookiestar28/ComfyUI-LongCat-Avatarhttps://github.com/rookiestar28/ComfyUI-LongCat-Avatar36MIT2026-062026-08LongCat Video Avatar 1.5 的 ComfyUI 自定义节点(音频驱动人物视频);另有 macOS-MLX 分支给 Apple Silicon部分(ComfyUI 节点,非服务化流式)
Kedreamix/Linly-Talker-Streamhttps://github.com/Kedreamix/Linly-Talker-Stream133Apache-2.02026-022026-02实时流式对话数字人系统:全双工、低延迟、实时交互✅ 实时流式对话(自述 full-duplex / low-latency / real-time)
taochangle/LingCasthttps://github.com/taochangle/LingCast11未标注2026-082026-08端到端 AI 数字人直播平台:LivePortrait 底视频 → Edge-TTS + Wav2Lip 口型 → DeepSeek 实时回复 → SRS 推流 + Next.js 观看端;Go(Gin)/MariaDB/Redis/RustFS/Docker + Python AI worker(macOS MPS / Linux CUDA / AMD ROCm)✅ 直播流式(SRS 推流链路)
avaturn-live/avtr-1https://github.com/avaturn-live/avtr-1469⚠️ 非开源许可:模型权重走 Community License(有营收门槛,超过需单独商业协议);renderer 与 live backend/streamer 均为 PolyForm Noncommercial License 1.0.0(仅非商用)(README「License」节)2026-052026-05AVTR-1:flow-matching 自回归模型做实时对话——输入一张肖像 + 双路音频,同时渲染说话口型 + 主动聆听,单卡 25 fps;含权重 / 推理代码 / 交互式流式 demo(TensorRT 加速,技术报告与生产后端标注 Coming soon)。要求 NVIDIA GPU(Ampere 及以上推荐)+ CUDA 12.x + TensorRT 10.x。模型卡 HF avaturn-live/avtr-1,项目页 https://avaturn-live.github.io/avtr-1-projectpage/✅ 流式实时——按 5 帧 chunk(25fps = 200ms/块)出图,README 给出逐卡延迟表:L40 84ms(2.4×) / A100 91ms(2.2×) / RTX 4060 Ti 166ms(1.2×) / RTX 3070 181ms(1.1×) / L4 202ms(0.99×) / RTX 3060 Ti 206ms(0.97×) / RTX 4060 232ms(0.86×) → 16GB 消费卡(4060 Ti 级)已能实时。⚠️ 商用需另谈授权
wpydcr/NanoAvatarhttps://github.com/wpydcr/NanoAvatar15代码 MIT / 口型权重 CC BY-NC 4.0(README「License」节明确分开说明)2026-092026-09 上周端侧 talking avatar:高通骁龙 8 Gen 3 上 39~41 FPS / 首帧 103~115ms / 显存仅 ~700–834 MiB;骁龙 8 Gen 1 仍有 18 FPS;RTX 4090 CUDA 全精度 224 FPS、量化 333 FPS。支持流式:"start speaking as audio arrives … starts speaking in about 0.3 seconds"。发布 Android APK(Full/Lite)+ HF 权重 wpydcr/NanoAvatar✅ 端侧/流式实时(显存需求全场最低,~0.8GB)— 佐证:HF 论坛帖
MrrZed0/Stream-Avatars-Customhttps://github.com/MrrZed0/Stream-Avatars-Custom1未标注2026-032026-03名字直指 "Stream Avatars" 的自定义版;仓库过小,描述未查到未查到
google/GNMhttps://github.com/google/GNM1.5kApache-2.02026-032026-09 上周五Google Generative aNthropometric Model:参数化统计人体/头部模型(GNM Head),NumPy/JAX/PyTorch/TF 后端,可商用;技术报告 2026-07 arXiv 2607.23687❌ 非 talking head(是头部几何模型,可作为下游资产)
NVlabs/SOMA-Xhttps://github.com/NVlabs/SOMA-X788Apache-2.02026-032026-09 上周四NVIDIA 参数化人体+手部可微表示(统一 SOMA 拓扑/骨骼,Warp 加速蒙皮拟合),兼容 MHR/Anny/SMPL/MANO;arXiv 2603.16858❌ 非 talking head(身体/手部几何与绑定)
k2-fsa/OmniVoicehttps://github.com/k2-fsa/OmniVoice14kApache-2.02026-032026-08新一代 TTS/语音模型(音频侧,非口型/人像)❌ 非 avatar(但是数字人语音前端候选)

(B) 已知项目在 2026 年的重大更新(供交叉参考)

仓库⭐License新建最后提交备注
Soul-AILab/SoulX-FlashHead1.1kApache-2.02026-022026-05本身就是 2026-02 新建:1.3B,单卡 RTX 4090 达 96 FPS 流式人像视频,无限时长。ithome 报道 · 极客公园
Soul-AILab/SoulX-FlashTalk1.5kApache-2.02025-122026-0714B,0.87s 启动延迟 / 32 FPS(8×H800)
lipku/LiveTalking9.6kApache-2.02023-122026-09仍在维护(已知名单)
HumanAIGC-Engineering/OpenAvatarChat3.8kApache-2.02025-022026-07已知名单,2026 仍活跃
modstart-lib/aigcpanel5.6kApache-2.02024-102026-09-22(今天)已知名单,2026 仍活跃
duixcom/Duix-Mobile8.3kGitHub 无法识别2024-052026-08已知名单,2026 仍活跃
duixcom/Duix-Avatar16kGitHub 无法识别2024-122026-04已知名单,2026 仍活跃
met4citizen/TalkingHead1.6kMIT2023-082026-06浏览器端 3D 说话头(已知名单)
jdh-algo/JoyVASA878MIT2024-112026-04京东音频驱动肖像动画(实时)
dlp3d-ai/dlp3d.ai358MIT2025-102026-05实时 3D 数字人(口型+动作)
Project-N-E-K-O/N.E.K.O3kApache-2.02025-062026-09-22(今天)3D AI 伙伴/数字人

(C) 对 16GB sm_120 单卡的选型提示(由本次调研直接推出)

  1. 显存最省 / sm_120 风险最低:NanoAvatar —— 权重显存 ~0.7–0.83 GB,且已在 RTX 4090 CUDA 上跑到 224–333 FPS,纯 CUDA 推理、不依赖 SageAttention/flash-attn;16GB 卡绰绰有余。注意口型权重是 CC BY-NC 4.0(非商用)。
  2. 最安全(3D 路线):SentiAvatar —— 不渲染像素、不碰 diffusers/SageAttention/FA,只跑 vLLM(Qwen2-0.5B) + Mask Transformer + RVQVAE,权重合计仅约 2.9GB(1.1GB + 276MB + 754MB + 50MB + 361MB + 1.5MB + 434MB,逐项见 README)。代价:输出是 BVH/UE 动作,需要自备渲染器(UE/Blender)。
  3. 流式视频里最有「消费卡可跑」证据的:AVTR-1(README 自附逐卡延迟表,4060 Ti 1.2× 实时);AvatarForcing(1.3B student + 34 ms/frame)与 SoulX-FlashHead(1.3B,4090 96FPS)也是小模型少步方案。但若实现依赖 flash-attn(sm_120 无预编译 wheel,见 1.3)或 SageAttention(见 1.2),务必先按 1.2(d) 换 wheel 或退回 SDPA。
  4. Super Star 的坑:其语音前端是 Qwen3-Omni-30B-A3B,对 16GB 显存不友好(README 建议 --gpu-memory-utilization 0.9 用 vLLM 部署),数字人单卡选型时需把它算进总账;动作侧(MotionGPT/RVQ-VAE)则很轻。
  5. 趋势判断:2026 年的实时数字人热点已明显从「LiveTalking/Wav2Lip 式拼装管线」转向 「少步流式扩散 talking head」(AvatarForcing one-step sliding-window、AVTR-1 flow-matching autoregressive、SoulX-FlashHead 96FPS@4090)与端侧(NanoAvatar)。

最终来源清单

SageAttention

  1. https://github.com/thu-ml/SageAttention/issues/391 (RTX 5060 Ti + H3 GPU lost;PyPI 只有 1.0.6)
  2. https://github.com/thu-ml/SageAttention/issues/357 (5060 Ti misaligned address;mengqin fork 修复)
  3. https://github.com/thu-ml/SageAttention/issues/291 (wgmma … not supported on .target 'sm_120')
  4. https://github.com/thu-ml/SageAttention/issues/303 (5060 Ti;_fused_attention False;setup.py 少逗号)
  5. https://github.com/thu-ml/SageAttention/issues/388 (>160k token 静默噪声)
  6. https://github.com/thu-ml/SageAttention/issues/392 (CUDA Graph 重放静默 90% 偏差;sm120→_qattn_sm89)
  7. https://github.com/thu-ml/SageAttention/issues/382 (TMA descriptor 初始化失败 ~74k token)
  8. https://github.com/thu-ml/SageAttention/issues/378 (sm_120 FP4 无收益)
  9. https://github.com/thu-ml/SageAttention/issues/330 (sm_121 NameError: num)
  10. https://github.com/thu-ml/SageAttention/issues/148 (RTX 50xx 安装失败)
  11. https://github.com/thu-ml/SageAttention/issues/107 (CUDA 12.8/torch 2.7 Blackwell 编译提问)
  12. https://github.com/thu-ml/SageAttention/issues/237 · https://github.com/thu-ml/SageAttention/issues/248 (Blackwell 家族支持请求)
  13. https://github.com/thu-ml/SageAttention/blob/main/README.md · https://github.com/thu-ml/SageAttention/blob/main/sageattention3_blackwell/README.md
  14. https://github.com/woct0rdho/SageAttention (预编译 sm120 wheel,v2.2.0-windows.post6)
  15. https://github.com/kijai/ComfyUI-KJNodes/issues/721 (H3 mem-eff SageAttention Patch 报错 + 修复步骤)
  16. https://github.com/kijai/ComfyUI-KJNodes/pull/729 (H3 sage sm90 分支 pad V 修复)
  17. https://github.com/mobcat40/sageattention-blackwell (第三方预编译 wheel + torch 2.11 头文件 patch)

FlashAttention / Triton

  1. https://github.com/Dao-AILab/flash-attention/issues/2307 (FA4 SM120 总状态)
  2. https://github.com/Dao-AILab/flash-attention/issues/1638 (FA2 sm_120 no kernel image)
  3. https://github.com/Dao-AILab/flash-attention/issues/2860 (FA4 sm120 varlen illegal memory access)
  4. https://github.com/Dao-AILab/flash-attention/issues/2832 (Windows + CUDA 13.4 + sm_121 构建失败)
  5. https://github.com/Dao-AILab/flash-attention/issues/1810 · https://github.com/Dao-AILab/flash-attention/issues/2444
  6. https://github.com/Dao-AILab/flash-attention/pull/2499 · /2634 · /2758 · /2771 · /2593 · /2862
  7. https://github.com/triton-lang/triton/pull/9734 · https://github.com/triton-lang/triton/pull/9755
  8. https://github.com/triton-lang/triton-windows
  9. https://github.com/sgl-project/sglang/pull/15776 · https://github.com/vllm-project/vllm/pull/47218

PyTorch / CUDA / xformers / ORT

  1. https://pytorch.org/blog/pytorch-2.7/ (2.7 = 首个 Blackwell 稳定支持 + cu128 + Triton 3.3)
  2. https://github.com/pytorch/pytorch/issues/159207 · /164342 (sm_120 官方支持请求)
  3. https://github.com/pytorch/pytorch/issues/172807 (sm_120a gencode 被剥 + PR #172721)
  4. https://github.com/pytorch/pytorch/issues/181379 (cuDNN SDPA head_dim=128 上限)
  5. https://github.com/pytorch/pytorch/issues/198126 · /191433 · /176426 · /186220 · /160838 · /178038
  6. https://pypi.org/pypi/torch/json · https://download.pytorch.org/whl/cu128/torch/ · .../cu129/torch/ · .../cu130/torch/
  7. https://github.com/facebookresearch/xformers/pull/1254 (Blackwell support,已 merged)
  8. https://github.com/microsoft/onnxruntime/issues/27621 (CUDA EP 线程静默死锁)
  9. https://github.com/Natfii/onnxruntime-gpu-blackwell · https://onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider
  10. https://blog.csdn.net/mrdeam/article/details/163705413 (ORT 1.26.0 Windows CUDA 12.9 sm_120 构建失败)

2026 仓库 / 社区

  1. https://github.com/SentiAvatar/SentiAvatar · https://arxiv.org/abs/2604.02908
  2. https://github.com/KlingAIResearch/AvatarForcing · https://arxiv.org/abs/2603.14331 · https://huggingface.co/lycui/AvatarForcing
  3. https://github.com/PeterIverson/Super-Star · https://arxiv.org/abs/2608.24909 · https://super-star-2026.github.io/
  4. https://github.com/rookiestar28/ComfyUI-LongCat-Avatar
  5. https://github.com/Kedreamix/Linly-Talker-Stream
  6. https://github.com/taochangle/LingCast
  7. https://github.com/avaturn-live/avtr-1 · https://huggingface.co/avaturn-live/avtr-1
  8. https://github.com/wpydcr/NanoAvatar · https://discuss.huggingface.co/t/nanoavatar-lifelike-talking-avatars-on-older-android-phones-no-cloud-gpu/180349
  9. https://github.com/google/GNM · https://arxiv.org/abs/2607.23687 · https://github.com/NVlabs/SOMA-X · https://arxiv.org/abs/2603.16858
  10. https://github.com/Soul-AILab/SoulX-FlashHead · https://m.ithome.com/html/921697.htm · https://w.geekpark.net/news/360311
  11. https://github.com/Soul-AILab/SoulX-FlashTalk
  12. https://github.com/topics/talking-head · https://github.com/topics/digital-human · https://github.com/topics/avatar-generation · https://github.com/topics/virtual-human · https://github.com/topics/talking-face · https://github.com/topics/lip-sync · https://github.com/trending?since=monthly
  13. https://img.shields.io/github/{stars,license,last-commit,created-at}/<owner>/<repo>.json (星数/许可/提交/建仓时间判定口径)

调研局限(须知)

下载此文件