调研日期:2026-09-22。目标硬件:RTX 5060 Ti 16GB / sm_120(消费级 Blackwell)/ 32GB RAM / Ubuntu 24.04。 所有结论均附可点击来源。查不到的写「未查到」,不做推测。 抓取方式:
curl直取 README / GitHub HTML / shields.io 徽章;github.com 的 HTML 在调研后段一度不可达(返回 000),因此少数条目只保留了已验证的徽章值。
这是一切坑的根因。sm_120(GeForce RTX 50 系)用的是 Ampere 时代的 mma.sync.aligned.m16n8k16(MMAv2),没有 data-center Blackwell(sm_100/103)才有的 tcgen05、WGMMA、Tensor Memory。
实践含义(对 16GB 5060 Ti):任何「只在 sm_90a / sm_100a 上编译」的 attention 内核(FA3、FA4、SageAttention3 FP4)在 sm_120 上要么编译失败、要么静默走退化路径。
getattr(sageattention, "_fused_attention", None) is not None → False,torch.cuda.get_device_capability() → (12, 0),相对 torch attention 几乎无加速。thu-ml/SageAttention#303「Does SageAttention 2.2 Support sm120 (SM 12.0) Blackwell?」SUPPORTED_ARCHS = {"8.0", "8.6", "8.9", "9.0", "10.0" "12.0"} —— "10.0" "12.0" 少了一个逗号,字符串被隐式拼接,导致 12.0 实际未被当作受支持架构。修复后提问者测到约 6% 加速。sageattn 在 (12,0) 上 dispatch 到 "sm120" 分支,而该分支实际调用 _qattn_sm89 扩展(sm89_compile.*),只是用 -gencode arch=compute_120a,code=sm_120a 把 SM89 内核编译成了 sm_120a 代码。thu-ml/SageAttention#392| # | 现象 / 报错原文 | 环境 | 来源 |
|---|---|---|---|
| 1 | 驱动级 GPU lost:No devices were found / GPU is lost(需重启);~54s 进入稳态生成时崩。原生 PyTorch attention 在更高显存(15.3GB vs 13.0GB)与更高功耗下跑通,排除显存/功耗/温度归因 | RTX 5060 Ti 16GB + MiniMax H3 + _sageattn_int8_fp8_nhd(int8-QK / per-channel FP8-PV) | #391 |
| 2 | RuntimeError: CUDA error: misaligned address(sageattn3_blackwell 推理时) | RTX 5060 Ti + PyTorch 2.12.0.dev+cu130 | #357 |
| 3 | ptxas error: Instruction 'wgmma.fence' not supported on .target 'sm_120';wgmma.mma_async with integer types / with FP8 types 同样不被支持(因 TORCH_CUDA_ARCH_LIST="9.0;12.0" 而误用 sm90 设置编 sm120) | RTX 5090, CUDA 12.8, torch 2.8.0+cu128 | #291 |
| 4 | 静默噪声:序列长度 >~160k token 时 FP8-PV 内核静默输出纯噪声(无异常、无 warning),采样器全步正常跑完;同配置关掉 SageAttention(SDPA)全部正常。~154k 以下干净,167k 起坏 | RTX PRO 6000 Blackwell(sm_120) + MiniMax H3 | #388 |
| 5 | CUDA Graph 重放静默错:capture 一次、之后往静态 buffer 写新值重放,输出相对 eager 偏差约 90%,无 NaN(下游解码成纯静态/噪声视频) | RTX PRO 6000 Blackwell(sm_120) | #392 |
| 6 | Error: Failed to initialize the TMA descriptor 1 → CUDA context illegal instruction(~73,774 token,per_block_mean=True;改 False 则正常) | sageattn3==1.0.0 sm_120a, torch 2.13.0+cu130 | #382 |
| 7 | 编译崩溃 NameError: name 'num' is not defined(setup.py 的 if/elif 链不覆盖 12.1) | DGX Spark sm_121 | #330 |
| 8 | SageAttention3 FP4 在 sm_120 上几乎无收益:内核级仅 1.17×、端到端 0.98×(论文在 5090 上是 1.84×) | RTX PRO 5000 Blackwell(sm_120) | #378 |
| 9 | ComfyUI/KJNodes 侧报错原文:RuntimeError: sageattention is not new enough version or could not determine CUDA architecture, cannot apply MiniMax H3 Memory Efficient Sage Attention Patch. | ComfyUI + MiniMax H3 | kijai/ComfyUI-KJNodes#721 |
本机自测记录的
sageattention/quant.pyper_channel_fp8OOM 这一具体报错文本,未查到对应的公开 issue 来源(GitHub HTML 调研后段不可达,无法穷举 issue 列表),故不列为有来源的结论;其根因与上表 #1/#4/#5 同源(sm_120 上走 SM89 内核路径)。
sageattention3_blackwell):面向 Blackwell 的 FP4 路径,需 sm_120a;官方 README 明确建议精度敏感场景仍用 SageAttention2("since SageAttention2 is more accurate, we still recommend using SageAttention2 for precision-sensitive applications")。SageAttention README · sageattention3_blackwell/README.mdsageattention3_blackwell/setup.py 的 Windows 分支加 -allow-unsupported-compiler,pip install --no-build-isolation -e .;提问者在 RTX 5060 Ti (sm_120) 上从 misaligned address 变为可运行(但性能提升有限)。另一用户同法修好了 5070 Ti。来源:#357 评论(讨论见 PR #323)v2.2.0-windows.post6(1k⭐,Apache-2.0,最后提交 2026-09)。README 原文:"The latest wheels support GTX 16xx, RTX 20xx/30xx/40xx/50xx, A100, H100, AGX Orin (sm75/80/86/87/89/90/120)" 与 "CUDA kernels for sm80/89/90 are bundled in the wheels, and also sm120 for CUDA >= 12.8"。另有 ABI3 稳定 ABI,torch2.10.0andhigher 一档通吃 torch ≥ 2.10。注意:CUDA 12 与 13 不通用。triton-windows<3.8 + woct0rdho v2.2.0-windows.post6 wheel)见 #721。pv_accum_dtype 调稳(sageattn_qk_int8_pv_fp16_cuda 最不易 overflow);见 woct0rdho README「Use notes」。| 版本 | sm_120 状态 | 来源 |
|---|---|---|
| FA2(flash-attn 2.x) | 预编译 wheel 不含 sm_120,直接报 RuntimeError: CUDA error: no kernel image is available for execution on the device。提问者反馈代码里最高只到 8.9;维护者答复编译器已指示支持 120(即需自行源码编译 + TORCH_CUDA_ARCH_LIST 含 12.0) | Dao-AILab/flash-attention#1638 |
| FA3 | 依赖 Hopper WGMMA,在 sm_120 上不可用,需显式 guard 以防误走 Hopper 路径 | PR #2593「Guard Hopper FA3 path against SM120 devices」 · #1810「FA3 attention sinks blackwell sm120」 |
| FA4(CuTe DSL) | 设计上走 tcgen05/TMEM → 原生不支持消费级 Blackwell;2026 年社区在补 sm120 路径(flash_fwd_sm120.py),但被 paged KV + num_splits 卡住,vLLM 服务化不可用;维护者原话倾向「不做」 | #2307 · 补丁:#2499 · #2634 · #2758 · #2771 |
| FA4 sm120 实跑 bug | flash_attn_varlen_func 在 ≥4 个 varlen 段、真实负载显存布局下 illegal memory access(GeForce Blackwell,测试于 MiniMax-H3 混合注意力视频模型,flash-attn-4 4.0.0b26/b29) | #2860 |
| 其它 sm120/SM80 路径 bug | pack_gqa 的 crd2idx 拒绝嵌套坐标;cu_seqlens[num_batch+1] OOB 读 | #2444 · PR #2862 |
| Windows 编译 | CUDA 13.4 + MSVC 传统预处理器:fatal error C1189: MSVC/cl.exe with traditional preprocessor is used … pass /Zc:preprocessor;/std:c++17 与 torch 2.15 头文件(需 C++20)冲突;且 无 sm_121 gencode | #2832 |
仓库基线:Dao-AILab/flash-attention(25k⭐,BSD-3-Clause,最后提交 2026-09 上旬)。
结论:在 5060 Ti 上不要指望 flash-attn;用 PyTorch SDPA 或(若必须)SageAttention2 的 sm120 wheel。
ptxas error: Instruction 'wgmma.fence' not supported on .target 'sm_120'。#291sm_120a,导致 LLVM/ptxas 为不存在的 tensor memory 特性生成指令 → 运行时 SIGSEGV。修复内容:sm_arch_from_capability 不再对 ≥90 一律加 "a";sm_120 改走 Hopper pipeline 而非 datacenter Blackwell pipeline。triton-lang/triton#9734torch.dot 在 RTX 5060 Ti (Blackwell, sm_120) 上 SIGFPE(exit 136)。pytorch/pytorch#178038pip install torch==2.7.0 --index-url https://download.pytorch.org/whl/cu128。 来源:PyTorch 2.7 Release Blog(原文行:"PyTorch 2.7 introduces support for NVIDIA's new Blackwell GPU architecture and ships pre-built wheels for CUDA 12.8")RuntimeError: CUDA error: no kernel image is available for execution on the device,或自行 TORCH_CUDA_ARCH_LIST="sm_120" 源码编译。来源:pytorch/pytorch#159207「Add official support for CUDA sm_120」 · #164342「Official support for sm_120 in stable PyTorch builds」cu131 返回 403(尚未开放)。索引:https://download.pytorch.org/whl/cu128/torch/、.../cu129/torch/、.../cu130/torch/https://pypi.org/pypi/torch/json → info.version;已发布版本尾段:2.10.0 / 2.11.0 / 2.12.0 / 2.12.1 / 2.13.0 / 2.14.0)。来源:pypi.org/pypi/torch/jsonsm_120a 后缀被自动剥掉。TORCH_CUDA_ARCH_LIST 未设置时,_get_cuda_arch_flags() 只取 major.minor,生成 -gencode=...sm_120 而非 sm_120a,直接破坏 CUTLASS block-scale MMA / NVFP4。影响 CUTLASS #2800 / #2820、TransformerEngine #2255。修复 PR #172721。来源:pytorch/pytorch#172807
TORCH_CUDA_ARCH_LIST="12.0a" 建出 ['sm_120a'],torch.zeros(...).to(torch.float4_e2m1fn_x2) 仍报 RuntimeError: copy_() does not support casting Float4_e2m1fn_x2 to different types. → 消费级 Blackwell 的 FP4 在 PyTorch 里仍不完整。check_cudnn_tensor_shapes 没有把 sm_120 放进已白名单的 [sm_80, sm_121] 区间),head_dim=256 的模型(Gemma 3、Qwen3 ≥14B、Llama 3.1 70B)会被静默拒绝并 fallthrough 到 MATH backend(O(n²) fp32 softmax)→ 长上下文 OOM。来源:pytorch/pytorch#181379[scaled_mm][mxfp8] 只有 TN 布局可用,NT/NN 被 cuBLAS 启发式拒绝 — #198126torch._inductor BF16 autocast 融合 embeddings/Linear/RMSNorm 静默算错 — #191433tl.load() 就 segfault — #176426this version of xformers requires torch 2.7。triton-windows,且实测需 triton-windows<3.8 才能与 SageAttention wheel 配合(KJNodes #721)。项目主页:triton-lang/triton-windowstl.load() segfault)。_is_fa4_supported() — vllm-project/vllm#47218。InferenceSession.run() 在主线程正常,放进 threading.Thread 就永久挂起(无异常、无超时)——发生在 RTX 5060 (Blackwell, sm_120), Windows,用 CUDAExecutionProvider 跑 RTMPose(rtmlib)。提问者归因于 PTX JIT 前向兼容(sm_120 非原生支持)与 GIL + 线程在 Windows 上的冲突。该 issue 已因 30 天无活动被 stale 机器人关闭。 规避:不要在 Python 线程里包 ONNX CUDA 推理。来源:microsoft/onnxruntime#27621onnxruntime-gpu 1.24.1 + Blackwell sm_120 CUDA kernels — Natfii/onnxruntime-gpu-blackwell(0⭐,无 license 标注,最后提交 2026-02)⚠️ opset 相关:本次调研未查到「Blackwell 与特定 ONNX opset 版本冲突」的一手官方来源,故不列结论。
判定口径:用 shields.io created-at 判「是否 2026 新建」,last-commit 判「是否有 2026 活动」;星数取 shields.io stars(2026-09-22 实测)。已排除题目给定的已知名单(除注明「2026 有重大更新」)。
| 仓库 | 链接 | ⭐ | License | 新建 | 最后提交 | 一句话 | 实时/流式? |
|---|---|---|---|---|---|---|---|
| SentiAvatar/SentiAvatar | https://github.com/SentiAvatar/SentiAvatar | 453 | GitHub 无法识别(论文 CC BY-NC-SA 4.0,见 arXiv 2604.02908) | 2026-04 | 2026-04 | SentiPulse(+ 人大 GSAI)开源的交互式 3D 数字人框架:SuSuInterActs 数据集(21K clip/37h) + 运动基础模型 + plan-then-infill + RVQVAE/Face VQVAE,输出 BVH / UE JSON(不是像素视频) | ✅ 实时:README 原文 "Generates 6 seconds of motion in 0.3 seconds with unlimited multi-turn streaming" |
| KlingAIResearch/AvatarForcing | https://github.com/KlingAIResearch/AvatarForcing | 81 | GitHub 无法识别 | 2026-03 | 2026-05 | 快手 Kling 的一步流式 talking avatar:单参考图 + 音频 + 文本 → 视频;局部未来滑窗 + 异构噪声联合去噪,每步产出 1 个干净 block,恒定单步开销。学生模型 1.3B(Wan2.1-T2V-1.3B),论文实测 34 ms/frame、25 FPS、832×480 | ✅ 流式(one-step streaming diffusion) — arXiv 2603.14331 |
| PeterIverson/Super-Star | https://github.com/PeterIverson/Super-Star | 10 | Apache-2.0 | 2026-07 | 2026-08 | [ACM MM 2026] "Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans"。3D 数字人在线实时交互:Streaming Speech Response(Qwen3-Omni-30B-A3B,需 CUDA driver ≥ 12.4)+ Online Gesture Generator(动作 RVQ-VAE + MotionGPT,torchrun --nproc_per_node=1 可单卡训练/推理);2026-08-14 放出训练/推理/评测代码,论文 arXiv 2608.24909,项目页 https://super-star-2026.github.io/ | ✅ 流式实时交互(在线 pipeline,含 streaming speech + online gesture) |
| rookiestar28/ComfyUI-LongCat-Avatar | https://github.com/rookiestar28/ComfyUI-LongCat-Avatar | 36 | MIT | 2026-06 | 2026-08 | LongCat Video Avatar 1.5 的 ComfyUI 自定义节点(音频驱动人物视频);另有 macOS-MLX 分支给 Apple Silicon | 部分(ComfyUI 节点,非服务化流式) |
| Kedreamix/Linly-Talker-Stream | https://github.com/Kedreamix/Linly-Talker-Stream | 133 | Apache-2.0 | 2026-02 | 2026-02 | 实时流式对话数字人系统:全双工、低延迟、实时交互 | ✅ 实时流式对话(自述 full-duplex / low-latency / real-time) |
| taochangle/LingCast | https://github.com/taochangle/LingCast | 11 | 未标注 | 2026-08 | 2026-08 | 端到端 AI 数字人直播平台:LivePortrait 底视频 → Edge-TTS + Wav2Lip 口型 → DeepSeek 实时回复 → SRS 推流 + Next.js 观看端;Go(Gin)/MariaDB/Redis/RustFS/Docker + Python AI worker(macOS MPS / Linux CUDA / AMD ROCm) | ✅ 直播流式(SRS 推流链路) |
| avaturn-live/avtr-1 | https://github.com/avaturn-live/avtr-1 | 469 | ⚠️ 非开源许可:模型权重走 Community License(有营收门槛,超过需单独商业协议);renderer 与 live backend/streamer 均为 PolyForm Noncommercial License 1.0.0(仅非商用)(README「License」节) | 2026-05 | 2026-05 | AVTR-1:flow-matching 自回归模型做实时对话——输入一张肖像 + 双路音频,同时渲染说话口型 + 主动聆听,单卡 25 fps;含权重 / 推理代码 / 交互式流式 demo(TensorRT 加速,技术报告与生产后端标注 Coming soon)。要求 NVIDIA GPU(Ampere 及以上推荐)+ CUDA 12.x + TensorRT 10.x。模型卡 HF avaturn-live/avtr-1,项目页 https://avaturn-live.github.io/avtr-1-projectpage/ | ✅ 流式实时——按 5 帧 chunk(25fps = 200ms/块)出图,README 给出逐卡延迟表:L40 84ms(2.4×) / A100 91ms(2.2×) / RTX 4060 Ti 166ms(1.2×) / RTX 3070 181ms(1.1×) / L4 202ms(0.99×) / RTX 3060 Ti 206ms(0.97×) / RTX 4060 232ms(0.86×) → 16GB 消费卡(4060 Ti 级)已能实时。⚠️ 商用需另谈授权 |
| wpydcr/NanoAvatar | https://github.com/wpydcr/NanoAvatar | 15 | 代码 MIT / 口型权重 CC BY-NC 4.0(README「License」节明确分开说明) | 2026-09 | 2026-09 上周 | 端侧 talking avatar:高通骁龙 8 Gen 3 上 39~41 FPS / 首帧 103~115ms / 显存仅 ~700–834 MiB;骁龙 8 Gen 1 仍有 18 FPS;RTX 4090 CUDA 全精度 224 FPS、量化 333 FPS。支持流式:"start speaking as audio arrives … starts speaking in about 0.3 seconds"。发布 Android APK(Full/Lite)+ HF 权重 wpydcr/NanoAvatar | ✅ 端侧/流式实时(显存需求全场最低,~0.8GB)— 佐证:HF 论坛帖 |
| MrrZed0/Stream-Avatars-Custom | https://github.com/MrrZed0/Stream-Avatars-Custom | 1 | 未标注 | 2026-03 | 2026-03 | 名字直指 "Stream Avatars" 的自定义版;仓库过小,描述未查到 | 未查到 |
| google/GNM | https://github.com/google/GNM | 1.5k | Apache-2.0 | 2026-03 | 2026-09 上周五 | Google Generative aNthropometric Model:参数化统计人体/头部模型(GNM Head),NumPy/JAX/PyTorch/TF 后端,可商用;技术报告 2026-07 arXiv 2607.23687 | ❌ 非 talking head(是头部几何模型,可作为下游资产) |
| NVlabs/SOMA-X | https://github.com/NVlabs/SOMA-X | 788 | Apache-2.0 | 2026-03 | 2026-09 上周四 | NVIDIA 参数化人体+手部可微表示(统一 SOMA 拓扑/骨骼,Warp 加速蒙皮拟合),兼容 MHR/Anny/SMPL/MANO;arXiv 2603.16858 | ❌ 非 talking head(身体/手部几何与绑定) |
| k2-fsa/OmniVoice | https://github.com/k2-fsa/OmniVoice | 14k | Apache-2.0 | 2026-03 | 2026-08 | 新一代 TTS/语音模型(音频侧,非口型/人像) | ❌ 非 avatar(但是数字人语音前端候选) |
| 仓库 | ⭐ | License | 新建 | 最后提交 | 备注 |
|---|---|---|---|---|---|
| Soul-AILab/SoulX-FlashHead | 1.1k | Apache-2.0 | 2026-02 | 2026-05 | 本身就是 2026-02 新建:1.3B,单卡 RTX 4090 达 96 FPS 流式人像视频,无限时长。ithome 报道 · 极客公园 |
| Soul-AILab/SoulX-FlashTalk | 1.5k | Apache-2.0 | 2025-12 | 2026-07 | 14B,0.87s 启动延迟 / 32 FPS(8×H800) |
| lipku/LiveTalking | 9.6k | Apache-2.0 | 2023-12 | 2026-09 | 仍在维护(已知名单) |
| HumanAIGC-Engineering/OpenAvatarChat | 3.8k | Apache-2.0 | 2025-02 | 2026-07 | 已知名单,2026 仍活跃 |
| modstart-lib/aigcpanel | 5.6k | Apache-2.0 | 2024-10 | 2026-09-22(今天) | 已知名单,2026 仍活跃 |
| duixcom/Duix-Mobile | 8.3k | GitHub 无法识别 | 2024-05 | 2026-08 | 已知名单,2026 仍活跃 |
| duixcom/Duix-Avatar | 16k | GitHub 无法识别 | 2024-12 | 2026-04 | 已知名单,2026 仍活跃 |
| met4citizen/TalkingHead | 1.6k | MIT | 2023-08 | 2026-06 | 浏览器端 3D 说话头(已知名单) |
| jdh-algo/JoyVASA | 878 | MIT | 2024-11 | 2026-04 | 京东音频驱动肖像动画(实时) |
| dlp3d-ai/dlp3d.ai | 358 | MIT | 2025-10 | 2026-05 | 实时 3D 数字人(口型+动作) |
| Project-N-E-K-O/N.E.K.O | 3k | Apache-2.0 | 2025-06 | 2026-09-22(今天) | 3D AI 伙伴/数字人 |
--gpu-memory-utilization 0.9 用 vLLM 部署),数字人单卡选型时需把它算进总账;动作侧(MotionGPT/RVQ-VAE)则很轻。SageAttention
misaligned address;mengqin fork 修复)wgmma … not supported on .target 'sm_120')_fused_attention False;setup.py 少逗号)_qattn_sm89)NameError: num)v2.2.0-windows.post6)FlashAttention / Triton
no kernel image)PyTorch / CUDA / xformers / ORT
2026 仓库 / 社区
raw.githubusercontent.com 与 img.shields.io 全程可用,因此所有星数/许可/提交时间/建仓时间均为实测徽章值,README 正文亦为实测原文。仅 MrrZed0/Stream-Avatars-Custom 因仓库过小未取到正文描述(已标注「未查到」)。per_channel_fp8 OOM 的具体报错文本未找到公开 issue,未作为有来源结论收录。april、september、last friday),不含年份。在调研日 2026-09-22 的语境下,无年份的同名月份一律解读为 2026 年;晚于 9 月的月份(october/november/december)一定属于更早年份,本表中已据此排除(例如 december 2025、november 2022)。