# Cluster D：sm_120/Blackwell 兼容性坑 + 2026 新增实时数字人仓库

> 调研日期：**2026-09-22**。目标硬件：**RTX 5060 Ti 16GB / sm_120（消费级 Blackwell）/ 32GB RAM / Ubuntu 24.04**。
> 所有结论均附可点击来源。**查不到的写「未查到」，不做推测。**
> 抓取方式：`curl` 直取 README / GitHub HTML / shields.io 徽章；github.com 的 HTML 在调研后段一度不可达（返回 000），因此少数条目只保留了已验证的徽章值。

---

## Section 1：sm_120 / Blackwell 兼容性坑（视频扩散 & 数字人栈）

### 1.1 总纲：消费级 Blackwell 没有 tcgen05 / WGMMA / TMEM

这是**一切坑的根因**。sm_120（GeForce RTX 50 系）用的是 **Ampere 时代的 `mma.sync.aligned.m16n8k16`（MMAv2）**，**没有** data-center Blackwell（sm_100/103）才有的 tcgen05、WGMMA、Tensor Memory。
- 来源（FlashAttention 维护者在 issue 中的定性）：[Dao-AILab/flash-attention#2307「FA4 SM120 support?」](https://github.com/Dao-AILab/flash-attention/issues/2307) — 原文：*"SM120 uses SM80-era mma.sync.aligned.m16n8k16 MMA instructions (no tcgen05/WGMMA/TMEM)."*
- 来源（Triton 侧同结论）：[triton-lang/triton#9734](https://github.com/triton-lang/triton/pull/9734) — 原文：*"There is no sm_120a. Consumer Blackwell is just sm_120."* / *"Consumer Blackwell uses MMAv2 … It has no tensor memory."*
- ⚠️ 该 Triton PR 后被 **revert**（arch 字符串结论本身被纠正，但 MMAv2 结论仍成立）：见 [#2307 中的 CORRECTION 评论](https://github.com/Dao-AILab/flash-attention/issues/2307) 指向 [triton-lang/triton#9755](https://github.com/triton-lang/triton/pull/9755)。

**实践含义（对 16GB 5060 Ti）**：任何「只在 sm_90a / sm_100a 上编译」的 attention 内核（FA3、FA4、SageAttention3 FP4）在 sm_120 上要么编译失败、要么静默走退化路径。

---

### 1.2 SageAttention on sm_120（坑最多，与本机 H3 掉总线问题直接相关）

#### (a) 官方 PyPI 包没有 sm_120 内核 → 静默回退 Triton JIT / 无加速
- issue #303（提问者就是 **RTX 5060 Ti 16GB**）：自建 wheel 后打印 `getattr(sageattention, "_fused_attention", None) is not None` → **`False`**，`torch.cuda.get_device_capability()` → `(12, 0)`，**相对 torch attention 几乎无加速**。[thu-ml/SageAttention#303「Does SageAttention 2.2 Support sm120 (SM 12.0) Blackwell?」](https://github.com/thu-ml/SageAttention/issues/303)
- 同 issue 中发现的 **setup.py 低级 bug**：`SUPPORTED_ARCHS = {"8.0", "8.6", "8.9", "9.0", "10.0" "12.0"}` —— `"10.0" "12.0"` **少了一个逗号**，字符串被隐式拼接，导致 12.0 实际未被当作受支持架构。修复后提问者测到约 **6%** 加速。
- issue #391 明确指出：**PyPI 上的 sageattention 只发布了 1.0.6**，要用 2.2.0 必须源码编译。[thu-ml/SageAttention#391](https://github.com/thu-ml/SageAttention/issues/391)
- 官方 README 只承认「编译代码支持 RTX5090」，并没有说预编译包带 sm_120 内核。[thu-ml/SageAttention README](https://github.com/thu-ml/SageAttention/blob/main/README.md)

#### (b) 关键内幕：sm_120 的 dispatch 实际调用的仍是 **SM89 内核**（不是原生 sm120 内核）
- issue #392 用 commit hash 定位：`sageattn` 在 `(12,0)` 上 dispatch 到 `"sm120"` 分支，而该分支**实际调用 `_qattn_sm89` 扩展**（`sm89_compile.*`），只是用 `-gencode arch=compute_120a,code=sm_120a` 把 SM89 内核编译成了 sm_120a 代码。[thu-ml/SageAttention#392](https://github.com/thu-ml/SageAttention/issues/392)
- **这解释了为什么 sm_120 上「能跑但会静默出错/驱动挂死」。**

#### (c) 具体报错清单（按严重度）

| # | 现象 / 报错原文 | 环境 | 来源 |
|---|---|---|---|
| 1 | **驱动级 GPU lost**：`No devices were found` / `GPU is lost`（需重启）；~54s 进入稳态生成时崩。原生 PyTorch attention 在**更高**显存（15.3GB vs 13.0GB）与更高功耗下**跑通**，排除显存/功耗/温度归因 | **RTX 5060 Ti 16GB** + MiniMax H3 + `_sageattn_int8_fp8_nhd`（int8-QK / per-channel FP8-PV） | [#391](https://github.com/thu-ml/SageAttention/issues/391) |
| 2 | `RuntimeError: CUDA error: misaligned address`（`sageattn3_blackwell` 推理时） | RTX 5060 Ti + PyTorch 2.12.0.dev+cu130 | [#357](https://github.com/thu-ml/SageAttention/issues/357) |
| 3 | `ptxas error: Instruction 'wgmma.fence' not supported on .target 'sm_120'`；`wgmma.mma_async with integer types` / `with FP8 types` 同样不被支持（因 `TORCH_CUDA_ARCH_LIST="9.0;12.0"` 而误用 sm90 设置编 sm120） | RTX 5090, CUDA 12.8, torch 2.8.0+cu128 | [#291](https://github.com/thu-ml/SageAttention/issues/291) |
| 4 | **静默噪声**：序列长度 >~160k token 时 FP8-PV 内核**静默输出纯噪声**（无异常、无 warning），采样器全步正常跑完；同配置关掉 SageAttention（SDPA）全部正常。~154k 以下干净，167k 起坏 | RTX PRO 6000 Blackwell(sm_120) + MiniMax H3 | [#388](https://github.com/thu-ml/SageAttention/issues/388) |
| 5 | **CUDA Graph 重放静默错**：capture 一次、之后往静态 buffer 写新值重放，输出相对 eager 偏差约 **90%**，无 NaN（下游解码成纯静态/噪声视频） | RTX PRO 6000 Blackwell(sm_120) | [#392](https://github.com/thu-ml/SageAttention/issues/392) |
| 6 | `Error: Failed to initialize the TMA descriptor 1` → CUDA context illegal instruction（~73,774 token，`per_block_mean=True`；改 `False` 则正常） | `sageattn3==1.0.0` sm_120a, torch 2.13.0+cu130 | [#382](https://github.com/thu-ml/SageAttention/issues/382) |
| 7 | 编译崩溃 `NameError: name 'num' is not defined`（setup.py 的 if/elif 链不覆盖 `12.1`） | DGX Spark sm_121 | [#330](https://github.com/thu-ml/SageAttention/issues/330) |
| 8 | SageAttention3 FP4 在 sm_120 上几乎无收益：内核级仅 1.17×、端到端 **0.98×**（论文在 5090 上是 1.84×） | RTX PRO 5000 Blackwell(sm_120) | [#378](https://github.com/thu-ml/SageAttention/issues/378) |
| 9 | ComfyUI/KJNodes 侧报错原文：`RuntimeError: sageattention is not new enough version or could not determine CUDA architecture, cannot apply MiniMax H3 Memory Efficient Sage Attention Patch.` | ComfyUI + MiniMax H3 | [kijai/ComfyUI-KJNodes#721](https://github.com/kijai/ComfyUI-KJNodes/issues/721) |

> 本机自测记录的 `sageattention/quant.py` **`per_channel_fp8` OOM** 这一具体报错文本，**未查到对应的公开 issue 来源**（GitHub HTML 调研后段不可达，无法穷举 issue 列表），故不列为有来源的结论；其根因与上表 #1/#4/#5 同源（sm_120 上走 SM89 内核路径）。

#### (d) 有 sm_120 内核的版本 / 已知修复

- **SageAttention 2.1 / 2.2**：走 SM89 内核 + sm_120a 编译（见 1.2b），「能跑」但精度/稳定性有风险。
- **SageAttention 3（`sageattention3_blackwell`）**：面向 Blackwell 的 FP4 路径，需 `sm_120a`；**官方 README 明确建议精度敏感场景仍用 SageAttention2**（*"since SageAttention2 is more accurate, we still recommend using SageAttention2 for precision-sensitive applications"*）。[SageAttention README](https://github.com/thu-ml/SageAttention/blob/main/README.md) · [sageattention3_blackwell/README.md](https://github.com/thu-ml/SageAttention/blob/main/sageattention3_blackwell/README.md)
- **✅ 已验证可用的修复路径（源码编译）**：用 [mengqin/SageAttention](https://github.com/mengqin/SageAttention) fork，在 `sageattention3_blackwell/setup.py` 的 Windows 分支加 `-allow-unsupported-compiler`，`pip install --no-build-isolation -e .`；提问者在 **RTX 5060 Ti (sm_120)** 上从 `misaligned address` 变为可运行（但性能提升有限）。另一用户同法修好了 **5070 Ti**。来源：[#357 评论](https://github.com/thu-ml/SageAttention/issues/357)（讨论见 [PR #323](https://github.com/thu-ml/SageAttention/pull/323)）
- **✅ 预编译 wheel（最省事）**：[woct0rdho/SageAttention](https://github.com/woct0rdho/SageAttention) fork，最新 release **`v2.2.0-windows.post6`**（1k⭐，Apache-2.0，最后提交 2026-09）。README 原文：*"The latest wheels support GTX 16xx, RTX 20xx/30xx/40xx/50xx, A100, H100, AGX Orin (sm75/80/86/87/89/90/120)"* 与 *"CUDA kernels for sm80/89/90 are bundled in the wheels, and also **sm120 for CUDA >= 12.8**"*。另有 ABI3 稳定 ABI，`torch2.10.0andhigher` 一档通吃 torch ≥ 2.10。注意：**CUDA 12 与 13 不通用**。
- ComfyUI 用户现在**可以不装第三方 wheel**，直接用 **comfy-kitchen** 里的 SageAttention（同上 README）；KJNodes 用户的实操修复帖（装 `triton-windows<3.8` + woct0rdho `v2.2.0-windows.post6` wheel）见 [#721](https://github.com/kijai/ComfyUI-KJNodes/issues/721)。
- 相关：KJNodes 针对 H3 长提示词崩溃的修复 [PR #729「pad V to CTA_K=128 in H3 mem-eff sage sm90 branch」](https://github.com/kijai/ComfyUI-KJNodes/pull/729)。

#### (e) 缓解建议（工程结论）
1. **在 5060 Ti 上，长序列视频扩散（H3/Wan 等）优先关掉 SageAttention**（用 PyTorch SDPA），#388/#391 都证明同一工作流 SDPA 正常。
2. 如必须开：用 woct0rdho **post6 / CUDA≥12.8** 的 sm120 wheel，并把 `pv_accum_dtype` 调稳（`sageattn_qk_int8_pv_fp16_cuda` 最不易 overflow）；见 [woct0rdho README「Use notes」](https://github.com/woct0rdho/SageAttention/blob/main/README.md)。
3. **不要用 SageAttention3 FP4 换速度**（sm_120 上无收益且更脆）。

---

### 1.3 FlashAttention on sm_120

| 版本 | sm_120 状态 | 来源 |
|---|---|---|
| **FA2（flash-attn 2.x）** | 预编译 wheel **不含 sm_120**，直接报 `RuntimeError: CUDA error: no kernel image is available for execution on the device`。提问者反馈代码里最高只到 8.9；维护者答复编译器已指示支持 120（即需自行源码编译 + `TORCH_CUDA_ARCH_LIST` 含 12.0） | [Dao-AILab/flash-attention#1638](https://github.com/Dao-AILab/flash-attention/issues/1638) |
| **FA3** | 依赖 Hopper WGMMA，在 sm_120 上不可用，需显式 guard 以防误走 Hopper 路径 | [PR #2593「Guard Hopper FA3 path against SM120 devices」](https://github.com/Dao-AILab/flash-attention/pull/2593) · [#1810「FA3 attention sinks blackwell sm120」](https://github.com/Dao-AILab/flash-attention/issues/1810) |
| **FA4（CuTe DSL）** | 设计上走 tcgen05/TMEM → **原生不支持消费级 Blackwell**；2026 年社区在补 sm120 路径（`flash_fwd_sm120.py`），但被 **paged KV + num_splits** 卡住，vLLM 服务化不可用；维护者原话倾向「不做」 | [#2307](https://github.com/Dao-AILab/flash-attention/issues/2307) · 补丁：[#2499](https://github.com/Dao-AILab/flash-attention/pull/2499) · [#2634](https://github.com/Dao-AILab/flash-attention/pull/2634) · [#2758](https://github.com/Dao-AILab/flash-attention/pull/2758) · [#2771](https://github.com/Dao-AILab/flash-attention/pull/2771) |
| **FA4 sm120 实跑 bug** | `flash_attn_varlen_func` 在 ≥4 个 varlen 段、真实负载显存布局下 **illegal memory access**（GeForce Blackwell，测试于 MiniMax-H3 混合注意力视频模型，flash-attn-4 4.0.0b26/b29） | [#2860](https://github.com/Dao-AILab/flash-attention/issues/2860) |
| **其它 sm120/SM80 路径 bug** | `pack_gqa` 的 `crd2idx` 拒绝嵌套坐标；`cu_seqlens[num_batch+1]` OOB 读 | [#2444](https://github.com/Dao-AILab/flash-attention/issues/2444) · [PR #2862](https://github.com/Dao-AILab/flash-attention/pull/2862) |
| **Windows 编译** | CUDA 13.4 + MSVC 传统预处理器：`fatal error C1189: MSVC/cl.exe with traditional preprocessor is used … pass /Zc:preprocessor`；`/std:c++17` 与 torch 2.15 头文件（需 C++20）冲突；且 **无 sm_121 gencode** | [#2832](https://github.com/Dao-AILab/flash-attention/issues/2832) |

仓库基线：[Dao-AILab/flash-attention](https://github.com/Dao-AILab/flash-attention)（25k⭐，BSD-3-Clause，最后提交 2026-09 上旬）。

**结论**：在 5060 Ti 上**不要指望 flash-attn**；用 PyTorch SDPA 或（若必须）SageAttention2 的 sm120 wheel。

---

### 1.4 wmma / wgmma 指令不兼容报告汇总

- **wgmma 不兼容（最直接的报错）**：见 1.2(c) #3，`ptxas error: Instruction 'wgmma.fence' not supported on .target 'sm_120'`。[#291](https://github.com/thu-ml/SageAttention/issues/291)
- 同一 issue 中社区澄清：**5090 系列根本不支持该指令，wgmma 只属于企业级 Blackwell**（原文：*"the 5090 series simply doesn't support that instruction. It is only for the enterprise blackwell GPU's."*）——同上链接。
- **Triton 曾把 sm_120 误判成 `sm_120a`**，导致 LLVM/ptxas 为不存在的 tensor memory 特性生成指令 → 运行时 SIGSEGV。修复内容：`sm_arch_from_capability` 不再对 ≥90 一律加 `"a"`；sm_120 改走 Hopper pipeline 而非 datacenter Blackwell pipeline。[triton-lang/triton#9734](https://github.com/triton-lang/triton/pull/9734)
- **PyTorch 侧**：`torch.dot` 在 **RTX 5060 Ti (Blackwell, sm_120)** 上 `SIGFPE`（exit 136）。[pytorch/pytorch#178038](https://github.com/pytorch/pytorch/issues/178038)

---

### 1.5 PyTorch / CUDA 版本要求（sm_120）

- **第一个正式支持 Blackwell 的稳定版是 PyTorch 2.7**，并发布 **CUDA 12.8 预编译 wheel**（Linux x86/arm64）；同时内置 **Triton 3.3**（*"adds support for the Blackwell architecture with torch.compile compatibility"*）。
  安装：`pip install torch==2.7.0 --index-url https://download.pytorch.org/whl/cu128`。
  来源：[PyTorch 2.7 Release Blog](https://pytorch.org/blog/pytorch-2.7/)（原文行：*"PyTorch 2.7 introduces support for NVIDIA's new Blackwell GPU architecture and ships pre-built wheels for CUDA 12.8"*）
- 在 2.7 之前，用户只能拿到 `RuntimeError: CUDA error: no kernel image is available for execution on the device`，或自行 `TORCH_CUDA_ARCH_LIST="sm_120"` 源码编译。来源：[pytorch/pytorch#159207「Add official support for CUDA sm_120」](https://github.com/pytorch/pytorch/issues/159207) · [#164342「Official support for sm_120 in stable PyTorch builds」](https://github.com/pytorch/pytorch/issues/164342)
- **cu128 / cu129 / cu130 三个 wheel 索引 2026-09-22 实测均可达**（HTTP 200）；`cu131` 返回 **403**（尚未开放）。索引：`https://download.pytorch.org/whl/cu128/torch/`、`.../cu129/torch/`、`.../cu130/torch/`
- **PyPI 上 torch 最新版 = 2.14.0**（实测 `https://pypi.org/pypi/torch/json` → `info.version`；已发布版本尾段：2.10.0 / 2.11.0 / 2.12.0 / 2.12.1 / 2.13.0 / 2.14.0）。来源：[pypi.org/pypi/torch/json](https://pypi.org/pypi/torch/json)
- **⚠️ 坑：`sm_120a` 后缀被自动剥掉**。`TORCH_CUDA_ARCH_LIST` 未设置时，`_get_cuda_arch_flags()` 只取 major.minor，生成 `-gencode=...sm_120` 而非 `sm_120a`，**直接破坏 CUTLASS block-scale MMA / NVFP4**。影响 CUTLASS [#2800](https://github.com/NVIDIA/cutlass/issues/2800) / [#2820](https://github.com/NVIDIA/cutlass/issues/2820)、TransformerEngine [#2255](https://github.com/NVIDIA/TransformerEngine/issues/2255)。修复 PR [#172721](https://github.com/pytorch/pytorch/pull/172721)。来源：[pytorch/pytorch#172807](https://github.com/pytorch/pytorch/issues/172807)
  - 该 issue 内 2026-05-03 的实测补充：即使显式 `TORCH_CUDA_ARCH_LIST="12.0a"` 建出 `['sm_120a']`，`torch.zeros(...).to(torch.float4_e2m1fn_x2)` 仍报 `RuntimeError: copy_() does not support casting Float4_e2m1fn_x2 to different types.` → **消费级 Blackwell 的 FP4 在 PyTorch 里仍不完整**。
- **cuDNN SDPA 在 sm_120 上 head_dim 上限卡在 128**（`check_cudnn_tensor_shapes` 没有把 sm_120 放进已白名单的 `[sm_80, sm_121]` 区间），head_dim=256 的模型（Gemma 3、Qwen3 ≥14B、Llama 3.1 70B）会被静默拒绝并 fallthrough 到 MATH backend（O(n²) fp32 softmax）→ **长上下文 OOM**。来源：[pytorch/pytorch#181379](https://github.com/pytorch/pytorch/issues/181379)
- **其它 sm_120 专属 PyTorch 坑**：
  - `[scaled_mm][mxfp8]` 只有 TN 布局可用，NT/NN 被 cuBLAS 启发式拒绝 — [#198126](https://github.com/pytorch/pytorch/issues/198126)
  - `torch._inductor` BF16 autocast 融合 embeddings/Linear/RMSNorm **静默算错** — [#191433](https://github.com/pytorch/pytorch/issues/191433)
  - Inductor Triton 在 sm_120 上**≥2 个 `tl.load()` 就 segfault** — [#176426](https://github.com/pytorch/pytorch/issues/176426)
  - 2.11.0+cu130 "duplicate template name" AssertionError（已 closed）— [#186220](https://github.com/pytorch/pytorch/issues/186220)
  - backward on 5090 比 4090 慢很多 — [#160838](https://github.com/pytorch/pytorch/issues/160838)

---

### 1.6 xformers / Triton 版本约束

- **xformers 直到 [PR #1254「Enabling Blackwell support」](https://github.com/facebookresearch/xformers/pull/1254) 才正式支持 Blackwell**：做法是「检查 CUDA ≥ 12.8 就启用 capability 120」（保守取 120 以覆盖全部 50 系）。该 PR **已 merged into main**（作者 loscrossos，维护者 danthe3rd 合并）。
  - PR 中用户实测反馈：拿到 wheel 后 **"no longer getting the error about the missing sm_120"**；但 **wheel 是为 torch 2.7 编译的**，torch 2.8.0.dev nightly 上会报 `this version of xformers requires torch 2.7`。
  - PR 中另一条关键反馈：**"The xformers wheel does not support the 50-series GPUs."**（官方预编译 wheel 当时不含 sm_120）
  - 仓库基线：[facebookresearch/xformers](https://github.com/facebookresearch/xformers)（11k⭐，GitHub 无法识别 license，最后提交 2026-09）
- **Triton**：
  - **Triton 3.3 随 PyTorch 2.7 首发支持 Blackwell**（[PyTorch 2.7 博客](https://pytorch.org/blog/pytorch-2.7/)）。
  - Windows 用户需用 **`triton-windows`**，且实测需 **`triton-windows<3.8`** 才能与 SageAttention wheel 配合（[KJNodes #721](https://github.com/kijai/ComfyUI-KJNodes/issues/721)）。项目主页：[triton-lang/triton-windows](https://github.com/triton-lang/triton-windows)
  - **sm_120 的 Triton 代码生成 bug**：[triton#9734](https://github.com/triton-lang/triton/pull/9734)（PTX codegen segfault；后被 revert，见 1.1）+ [pytorch#176426](https://github.com/pytorch/pytorch/issues/176426)（≥2 个 `tl.load()` segfault）。
  - 对照：SGLang 曾直接 **在 SM120 上 fall back 到 SDPA**，因为 diffusion 路径不稳 — [sgl-project/sglang#15776](https://github.com/sgl-project/sglang/pull/15776)；vLLM 亦需把 SM12x 加进 `_is_fa4_supported()` — [vllm-project/vllm#47218](https://github.com/vllm-project/vllm/pull/47218)。

---

### 1.7 ONNX Runtime / CUDA EP 在 Blackwell 上的坑

- **CUDA EP 静默死锁**：`InferenceSession.run()` 在主线程正常，**放进 `threading.Thread` 就永久挂起**（无异常、无超时）——发生在 **RTX 5060 (Blackwell, sm_120), Windows**，用 `CUDAExecutionProvider` 跑 RTMPose（rtmlib）。提问者归因于 **PTX JIT 前向兼容（sm_120 非原生支持）与 GIL + 线程在 Windows 上的冲突**。**该 issue 已因 30 天无活动被 stale 机器人关闭。** 规避：不要在 Python 线程里包 ONNX CUDA 推理。来源：[microsoft/onnxruntime#27621](https://github.com/microsoft/onnxruntime/issues/27621)
- **官方 wheel 不含 sm_120 内核**，社区自建：`onnxruntime-gpu 1.24.1` + Blackwell sm_120 CUDA kernels — [Natfii/onnxruntime-gpu-blackwell](https://github.com/Natfii/onnxruntime-gpu-blackwell)（0⭐，无 license 标注，最后提交 2026-02）
- **Windows + CUDA 12.9 + sm_120 上 ORT 1.26.0 构建失败**（第三方记录）：[CSDN 修复记录](https://blog.csdn.net/mrdeam/article/details/163705413)
- 官方 CUDA EP 文档（版本/依赖矩阵）：[onnxruntime.ai CUDA Execution Provider](https://onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider)
- 仓库基线：[microsoft/onnxruntime](https://github.com/microsoft/onnxruntime)（22k⭐，MIT，最后提交 2026-09-22）

> ⚠️ **opset 相关**：本次调研**未查到**「Blackwell 与特定 ONNX opset 版本冲突」的一手官方来源，故不列结论。

---

## Section 2：2026 年（1–9 月）新增 / 新活跃的实时流式数字人仓库

**判定口径**：用 shields.io `created-at` 判「是否 2026 新建」，`last-commit` 判「是否有 2026 活动」；星数取 shields.io `stars`（2026-09-22 实测）。**已排除**题目给定的已知名单（除注明「2026 有重大更新」）。

### (A) 2026 年**全新**仓库

| 仓库 | 链接 | ⭐ | License | 新建 | 最后提交 | 一句话 | 实时/流式？ |
|---|---|---|---|---|---|---|---|
| **SentiAvatar/SentiAvatar** | https://github.com/SentiAvatar/SentiAvatar | **453** | GitHub 无法识别（**论文 CC BY-NC-SA 4.0**，见 [arXiv 2604.02908](https://arxiv.org/abs/2604.02908)） | 2026-04 | 2026-04 | SentiPulse（+ 人大 GSAI）开源的**交互式 3D 数字人**框架：SuSuInterActs 数据集(21K clip/37h) + 运动基础模型 + plan-then-infill + RVQVAE/Face VQVAE，输出 **BVH / UE JSON**（**不是像素视频**） | **✅ 实时**：README 原文 *"Generates 6 seconds of motion in 0.3 seconds with unlimited multi-turn streaming"* |
| **KlingAIResearch/AvatarForcing** | https://github.com/KlingAIResearch/AvatarForcing | **81** | GitHub 无法识别 | 2026-03 | 2026-05 | 快手 Kling 的**一步流式 talking avatar**：单参考图 + 音频 + 文本 → 视频；局部未来滑窗 + 异构噪声联合去噪，每步产出 1 个干净 block，恒定单步开销。学生模型 **1.3B**（Wan2.1-T2V-1.3B），论文实测 **34 ms/frame**、25 FPS、832×480 | **✅ 流式（one-step streaming diffusion）** — [arXiv 2603.14331](https://arxiv.org/abs/2603.14331) |
| **PeterIverson/Super-Star** | https://github.com/PeterIverson/Super-Star | **10** | Apache-2.0 | 2026-07 | 2026-08 | **[ACM MM 2026]** "Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans"。**3D 数字人**在线实时交互：Streaming Speech Response（Qwen3-Omni-30B-A3B，**需 CUDA driver ≥ 12.4**）+ Online Gesture Generator（动作 RVQ-VAE + MotionGPT，`torchrun --nproc_per_node=1` 可单卡训练/推理）；2026-08-14 放出训练/推理/评测代码，论文 [arXiv 2608.24909](https://arxiv.org/abs/2608.24909)，项目页 https://super-star-2026.github.io/ | **✅ 流式实时交互**（在线 pipeline，含 streaming speech + online gesture） |
| **rookiestar28/ComfyUI-LongCat-Avatar** | https://github.com/rookiestar28/ComfyUI-LongCat-Avatar | **36** | MIT | 2026-06 | 2026-08 | **LongCat Video Avatar 1.5** 的 ComfyUI 自定义节点（音频驱动人物视频）；另有 macOS-MLX 分支给 Apple Silicon | 部分（ComfyUI 节点，非服务化流式） |
| **Kedreamix/Linly-Talker-Stream** | https://github.com/Kedreamix/Linly-Talker-Stream | **133** | Apache-2.0 | 2026-02 | 2026-02 | **实时流式对话数字人系统**：全双工、低延迟、实时交互 | **✅ 实时流式对话**（自述 full-duplex / low-latency / real-time） |
| **taochangle/LingCast** | https://github.com/taochangle/LingCast | **11** | 未标注 | 2026-08 | 2026-08 | 端到端 AI 数字人**直播**平台：LivePortrait 底视频 → Edge-TTS + Wav2Lip 口型 → DeepSeek 实时回复 → SRS 推流 + Next.js 观看端；Go(Gin)/MariaDB/Redis/RustFS/Docker + Python AI worker（macOS MPS / Linux CUDA / AMD ROCm） | **✅ 直播流式**（SRS 推流链路） |
| **avaturn-live/avtr-1** | https://github.com/avaturn-live/avtr-1 | **469** | ⚠️ **非开源许可**：模型权重走 **Community License**（**有营收门槛**，超过需单独商业协议）；renderer 与 live backend/streamer 均为 **PolyForm Noncommercial License 1.0.0（仅非商用）**（README「License」节） | 2026-05 | 2026-05 | **AVTR-1**：flow-matching 自回归模型做**实时对话**——输入一张肖像 + 双路音频，同时渲染**说话口型 + 主动聆听**，单卡 **25 fps**；含权重 / 推理代码 / 交互式流式 demo（TensorRT 加速，技术报告与生产后端标注 Coming soon）。要求 NVIDIA GPU（Ampere 及以上推荐）+ **CUDA 12.x + TensorRT 10.x**。模型卡 [HF avaturn-live/avtr-1](https://huggingface.co/avaturn-live/avtr-1)，项目页 https://avaturn-live.github.io/avtr-1-projectpage/ | **✅ 流式实时**——按 5 帧 chunk（25fps = 200ms/块）出图，README 给出**逐卡延迟表**：L40 84ms(2.4×) / A100 91ms(2.2×) / **RTX 4060 Ti 166ms(1.2×)** / RTX 3070 181ms(1.1×) / L4 202ms(0.99×) / RTX 3060 Ti 206ms(0.97×) / RTX 4060 232ms(0.86×) → **16GB 消费卡（4060 Ti 级）已能实时**。⚠️ **商用需另谈授权** |
| **wpydcr/NanoAvatar** | https://github.com/wpydcr/NanoAvatar | **15** | 代码 **MIT** / 口型权重 **CC BY-NC 4.0**（README「License」节明确分开说明） | 2026-09 | 2026-09 上周 | **端侧 talking avatar**：高通骁龙 8 Gen 3 上 **39~41 FPS / 首帧 103~115ms / 显存仅 ~700–834 MiB**；骁龙 8 Gen 1 仍有 18 FPS；**RTX 4090 CUDA 全精度 224 FPS、量化 333 FPS**。支持流式：*"start speaking as audio arrives … starts speaking in about 0.3 seconds"*。发布 Android APK（Full/Lite）+ [HF 权重 wpydcr/NanoAvatar](https://huggingface.co/wpydcr/NanoAvatar) | **✅ 端侧/流式实时**（**显存需求全场最低，~0.8GB**）— 佐证：[HF 论坛帖](https://discuss.huggingface.co/t/nanoavatar-lifelike-talking-avatars-on-older-android-phones-no-cloud-gpu/180349) |
| **MrrZed0/Stream-Avatars-Custom** | https://github.com/MrrZed0/Stream-Avatars-Custom | **1** | 未标注 | 2026-03 | 2026-03 | 名字直指 "Stream Avatars" 的自定义版；仓库过小，**描述未查到** | **未查到** |
| **google/GNM** | https://github.com/google/GNM | **1.5k** | Apache-2.0 | 2026-03 | 2026-09 上周五 | Google **Generative aNthropometric Model**：参数化统计人体/头部模型（GNM Head），NumPy/JAX/PyTorch/TF 后端，可商用；技术报告 2026-07 [arXiv 2607.23687](https://arxiv.org/abs/2607.23687) | ❌ 非 talking head（是头部**几何**模型，可作为下游资产） |
| **NVlabs/SOMA-X** | https://github.com/NVlabs/SOMA-X | **788** | Apache-2.0 | 2026-03 | 2026-09 上周四 | NVIDIA 参数化人体+手部**可微表示**（统一 SOMA 拓扑/骨骼，Warp 加速蒙皮拟合），兼容 MHR/Anny/SMPL/MANO；[arXiv 2603.16858](https://arxiv.org/abs/2603.16858) | ❌ 非 talking head（身体/手部几何与绑定） |
| **k2-fsa/OmniVoice** | https://github.com/k2-fsa/OmniVoice | **14k** | Apache-2.0 | 2026-03 | 2026-08 | 新一代 TTS/语音模型（**音频侧**，非口型/人像） | ❌ 非 avatar（但是数字人语音前端候选） |

### (B) 已知项目在 2026 年的重大更新（供交叉参考）

| 仓库 | ⭐ | License | 新建 | 最后提交 | 备注 |
|---|---|---|---|---|---|
| [Soul-AILab/SoulX-FlashHead](https://github.com/Soul-AILab/SoulX-FlashHead) | 1.1k | Apache-2.0 | 2026-02 | 2026-05 | **本身就是 2026-02 新建**：1.3B，单卡 **RTX 4090 达 96 FPS** 流式人像视频，无限时长。[ithome 报道](https://m.ithome.com/html/921697.htm) · [极客公园](https://w.geekpark.net/news/360311) |
| [Soul-AILab/SoulX-FlashTalk](https://github.com/Soul-AILab/SoulX-FlashTalk) | 1.5k | Apache-2.0 | 2025-12 | **2026-07** | 14B，**0.87s 启动延迟 / 32 FPS**（8×H800） |
| [lipku/LiveTalking](https://github.com/lipku/LiveTalking) | 9.6k | Apache-2.0 | 2023-12 | **2026-09** | 仍在维护（已知名单） |
| [HumanAIGC-Engineering/OpenAvatarChat](https://github.com/HumanAIGC-Engineering/OpenAvatarChat) | 3.8k | Apache-2.0 | 2025-02 | **2026-07** | 已知名单，2026 仍活跃 |
| [modstart-lib/aigcpanel](https://github.com/modstart-lib/aigcpanel) | 5.6k | Apache-2.0 | 2024-10 | **2026-09-22（今天）** | 已知名单，2026 仍活跃 |
| [duixcom/Duix-Mobile](https://github.com/duixcom/Duix-Mobile) | 8.3k | GitHub 无法识别 | 2024-05 | **2026-08** | 已知名单，2026 仍活跃 |
| [duixcom/Duix-Avatar](https://github.com/duixcom/Duix-Avatar) | 16k | GitHub 无法识别 | 2024-12 | **2026-04** | 已知名单，2026 仍活跃 |
| [met4citizen/TalkingHead](https://github.com/met4citizen/TalkingHead) | 1.6k | MIT | 2023-08 | **2026-06** | 浏览器端 3D 说话头（已知名单） |
| [jdh-algo/JoyVASA](https://github.com/jdh-algo/JoyVASA) | 878 | MIT | 2024-11 | **2026-04** | 京东音频驱动肖像动画（**实时**） |
| [dlp3d-ai/dlp3d.ai](https://github.com/dlp3d-ai/dlp3d.ai) | 358 | MIT | 2025-10 | **2026-05** | 实时 3D 数字人（口型+动作） |
| [Project-N-E-K-O/N.E.K.O](https://github.com/Project-N-E-K-O/N.E.K.O) | 3k | Apache-2.0 | 2025-06 | **2026-09-22（今天）** | 3D AI 伙伴/数字人 |

### (C) 对 16GB sm_120 单卡的选型提示（由本次调研直接推出）

1. **显存最省 / sm_120 风险最低**：**NanoAvatar** —— 权重显存 ~**0.7–0.83 GB**，且已在 **RTX 4090 CUDA** 上跑到 224–333 FPS，纯 CUDA 推理、**不依赖 SageAttention/flash-attn**；16GB 卡绰绰有余。注意口型权重是 **CC BY-NC 4.0（非商用）**。
2. **最安全（3D 路线）**：**SentiAvatar** —— 不渲染像素、不碰 diffusers/SageAttention/FA，只跑 vLLM(Qwen2-0.5B) + Mask Transformer + RVQVAE，权重合计仅约 **2.9GB**（1.1GB + 276MB + 754MB + 50MB + 361MB + 1.5MB + 434MB，逐项见 README）。代价：输出是 **BVH/UE 动作**，需要自备渲染器（UE/Blender）。
3. **流式视频里最有「消费卡可跑」证据的**：**AVTR-1**（README 自附逐卡延迟表，4060 Ti 1.2× 实时）；**AvatarForcing**（1.3B student + 34 ms/frame）与 **SoulX-FlashHead**（1.3B，4090 96FPS）也是小模型少步方案。但若实现依赖 flash-attn（sm_120 无预编译 wheel，见 1.3）或 SageAttention（见 1.2），务必先按 1.2(d) 换 wheel 或退回 SDPA。
4. **Super Star 的坑**：其语音前端是 **Qwen3-Omni-30B-A3B**，对 16GB 显存不友好（README 建议 `--gpu-memory-utilization 0.9` 用 vLLM 部署），数字人单卡选型时需把它算进总账；动作侧（MotionGPT/RVQ-VAE）则很轻。
5. **趋势判断**：2026 年的实时数字人热点已明显从「LiveTalking/Wav2Lip 式拼装管线」转向 **「少步流式扩散 talking head」**（AvatarForcing one-step sliding-window、AVTR-1 flow-matching autoregressive、SoulX-FlashHead 96FPS@4090）与**端侧**（NanoAvatar）。

---

## 最终来源清单

**SageAttention**
1. https://github.com/thu-ml/SageAttention/issues/391 （RTX 5060 Ti + H3 GPU lost；PyPI 只有 1.0.6）
2. https://github.com/thu-ml/SageAttention/issues/357 （5060 Ti `misaligned address`；mengqin fork 修复）
3. https://github.com/thu-ml/SageAttention/issues/291 （`wgmma … not supported on .target 'sm_120'`）
4. https://github.com/thu-ml/SageAttention/issues/303 （5060 Ti；`_fused_attention` False；setup.py 少逗号）
5. https://github.com/thu-ml/SageAttention/issues/388 （>160k token 静默噪声）
6. https://github.com/thu-ml/SageAttention/issues/392 （CUDA Graph 重放静默 90% 偏差；sm120→`_qattn_sm89`）
7. https://github.com/thu-ml/SageAttention/issues/382 （TMA descriptor 初始化失败 ~74k token）
8. https://github.com/thu-ml/SageAttention/issues/378 （sm_120 FP4 无收益）
9. https://github.com/thu-ml/SageAttention/issues/330 （sm_121 `NameError: num`）
10. https://github.com/thu-ml/SageAttention/issues/148 （RTX 50xx 安装失败）
11. https://github.com/thu-ml/SageAttention/issues/107 （CUDA 12.8/torch 2.7 Blackwell 编译提问）
12. https://github.com/thu-ml/SageAttention/issues/237 · https://github.com/thu-ml/SageAttention/issues/248 （Blackwell 家族支持请求）
13. https://github.com/thu-ml/SageAttention/blob/main/README.md · https://github.com/thu-ml/SageAttention/blob/main/sageattention3_blackwell/README.md
14. https://github.com/woct0rdho/SageAttention （预编译 sm120 wheel，`v2.2.0-windows.post6`）
15. https://github.com/kijai/ComfyUI-KJNodes/issues/721 （H3 mem-eff SageAttention Patch 报错 + 修复步骤）
16. https://github.com/kijai/ComfyUI-KJNodes/pull/729 （H3 sage sm90 分支 pad V 修复）
17. https://github.com/mobcat40/sageattention-blackwell （第三方预编译 wheel + torch 2.11 头文件 patch）

**FlashAttention / Triton**
18. https://github.com/Dao-AILab/flash-attention/issues/2307 （FA4 SM120 总状态）
19. https://github.com/Dao-AILab/flash-attention/issues/1638 （FA2 sm_120 `no kernel image`）
20. https://github.com/Dao-AILab/flash-attention/issues/2860 （FA4 sm120 varlen illegal memory access）
21. https://github.com/Dao-AILab/flash-attention/issues/2832 （Windows + CUDA 13.4 + sm_121 构建失败）
22. https://github.com/Dao-AILab/flash-attention/issues/1810 · https://github.com/Dao-AILab/flash-attention/issues/2444
23. https://github.com/Dao-AILab/flash-attention/pull/2499 · /2634 · /2758 · /2771 · /2593 · /2862
24. https://github.com/triton-lang/triton/pull/9734 · https://github.com/triton-lang/triton/pull/9755
25. https://github.com/triton-lang/triton-windows
26. https://github.com/sgl-project/sglang/pull/15776 · https://github.com/vllm-project/vllm/pull/47218

**PyTorch / CUDA / xformers / ORT**
27. https://pytorch.org/blog/pytorch-2.7/ （2.7 = 首个 Blackwell 稳定支持 + cu128 + Triton 3.3）
28. https://github.com/pytorch/pytorch/issues/159207 · /164342 （sm_120 官方支持请求）
29. https://github.com/pytorch/pytorch/issues/172807 （sm_120a gencode 被剥 + PR #172721）
30. https://github.com/pytorch/pytorch/issues/181379 （cuDNN SDPA head_dim=128 上限）
31. https://github.com/pytorch/pytorch/issues/198126 · /191433 · /176426 · /186220 · /160838 · /178038
32. https://pypi.org/pypi/torch/json · https://download.pytorch.org/whl/cu128/torch/ · .../cu129/torch/ · .../cu130/torch/
33. https://github.com/facebookresearch/xformers/pull/1254 （Blackwell support，已 merged）
34. https://github.com/microsoft/onnxruntime/issues/27621 （CUDA EP 线程静默死锁）
35. https://github.com/Natfii/onnxruntime-gpu-blackwell · https://onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider
36. https://blog.csdn.net/mrdeam/article/details/163705413 （ORT 1.26.0 Windows CUDA 12.9 sm_120 构建失败）

**2026 仓库 / 社区**
37. https://github.com/SentiAvatar/SentiAvatar · https://arxiv.org/abs/2604.02908
38. https://github.com/KlingAIResearch/AvatarForcing · https://arxiv.org/abs/2603.14331 · https://huggingface.co/lycui/AvatarForcing
39. https://github.com/PeterIverson/Super-Star · https://arxiv.org/abs/2608.24909 · https://super-star-2026.github.io/
40. https://github.com/rookiestar28/ComfyUI-LongCat-Avatar
41. https://github.com/Kedreamix/Linly-Talker-Stream
42. https://github.com/taochangle/LingCast
43. https://github.com/avaturn-live/avtr-1 · https://huggingface.co/avaturn-live/avtr-1
44. https://github.com/wpydcr/NanoAvatar · https://discuss.huggingface.co/t/nanoavatar-lifelike-talking-avatars-on-older-android-phones-no-cloud-gpu/180349
45. https://github.com/google/GNM · https://arxiv.org/abs/2607.23687 · https://github.com/NVlabs/SOMA-X · https://arxiv.org/abs/2603.16858
46. https://github.com/Soul-AILab/SoulX-FlashHead · https://m.ithome.com/html/921697.htm · https://w.geekpark.net/news/360311
47. https://github.com/Soul-AILab/SoulX-FlashTalk
48. https://github.com/topics/talking-head · https://github.com/topics/digital-human · https://github.com/topics/avatar-generation · https://github.com/topics/virtual-human · https://github.com/topics/talking-face · https://github.com/topics/lip-sync · https://github.com/trending?since=monthly
49. https://img.shields.io/github/{stars,license,last-commit,created-at}/<owner>/<repo>.json （星数/许可/提交/建仓时间判定口径）

---

### 调研局限（须知）
- **github.com 的 HTML 页在本次调研后段不可达（curl 返回 000 / 超时）**，GitHub REST API 未认证额度也已耗尽（60/h）；但 `raw.githubusercontent.com` 与 `img.shields.io` 全程可用，因此所有星数/许可/提交时间/建仓时间**均为实测徽章值**，README 正文亦为实测原文。仅 **MrrZed0/Stream-Avatars-Custom** 因仓库过小未取到正文描述（已标注「未查到」）。
- **`per_channel_fp8` OOM** 的具体报错文本未找到公开 issue，未作为有来源结论收录。
- **ONNX Runtime 与 opset 版本的冲突**未查到一手来源。
- ⚠️ shields.io 的「last commit / created at」对**当年**只返回**月份单词**（如 `april`、`september`、`last friday`），不含年份。在调研日 2026-09-22 的语境下，无年份的同名月份一律解读为 **2026 年**；**晚于 9 月的月份（october/november/december）一定属于更早年份**，本表中已据此排除（例如 `december 2025`、`november 2022`）。
