findings_C2.md

C2 组核查结果:实时数字人(talking head)开源项目 — 3 仓深挖


0. 头条结论(先说最重要的三条)

  1. duixcom/Duix.Heygem 这个仓库在 2026 年已经不存在了 —— 它被官方重命名为 duixcom/Duix-Avatar,GitHub 对旧地址做 301 跳转。所以"Duix.Heygem 2026 状态"的正确答案是:主体已改名,改名后的 Duix-Avatar 是唯一活跃主线。
  2. Duix.Avatar 在 Blackwell / sm_120 上目前是坏的:2026-09-04 有人开了 Issue #624,报 CUDA error: no kernel image is available for execution on the device,根因指向 Docker 镜像里的 PyTorch 不含 CC 12.0 内核,该 Issue 至今 Open 且零回复。
  3. LatentSync 官方 1.6 的最低推理显存是 18GB,16GB 单卡不达官方门槛;但 2026-07-03 有人在 RTX 5070 Ti 16GB 上实测跑通了(峰值 15807MB),代价是约 27.7–41.1 秒墙钟时间换 1 秒输出 —— 约为实时的 1/28 ~ 1/41。

1. duixcom/Duix.Heygem(→ 已改名 duixcom/Duix-Avatar)

1.1 仓库身份与 2026 活跃度

项目结果证据
duixcom/Duix.Heygem301 跳转到 duixcom/Duix-Avatarcurl -sL -w '%{url_effective}' → 200 -> https://github.com/duixcom/Duix-Avatar;ungh.cc 对 duixcom/Duix.Heygem 查询返回 "id":907627874,"name":"Duix-Avatar"(同一个 repo id,证明是重命名而非新仓)
Stars / Forks15547 / 2642ungh.cc/repos/duixcom/Duix-Avatar,2026-09-22 读取
创建 / 最后 push创建 2024-12-24T03:00:26Z;最后 push 2026-04-21T07:06:36Z同上。→ 2026 年有提交(4 月),但到 9 月已约 5 个月无代码更新
改名证据(Release)v1.0.5(2025-08-15)发布说明:"The original project name "HeyGem" has now been officially changed to "Duix.Avatar".";v1.0.6(2025-09-28):"It's just a name change. If the version you're using has no issues, you can continue to use it."releases

1.2 duixcom org 下各仓 2026 状态(逐个核查,含"是否存在")

说明:任务要求的 curl https://api.github.com/orgs/duixcom/repos 因本机 IP 限流(core remaining=0,reset 13:52)无法执行;改用 ungh.cc 镜像逐个精确查询 + 落地页跳转探测,结论同样可证伪。

仓库状态(2026-09-22)Stars最后 push说明
duixcom/Duix-Avatar(原 Duix.Heygem)✅ 存在、2026 有更新155472026-04-21本组主目标的正身
duixcom/Duix.Mobile(原 Duix.mobile)✅ 存在、2026 活跃82522026-08-05移动端/嵌入式实时交互 SDK,非本地 GPU 服务。仓描述:"on-premise deployment and <1.5 s latency"
duixcom/Duix.Heygem.Android❌ 404,不存在——该仓从未存在或已删除;raw 的 main/master 均 404
duixcom/duix-skills⚠️ 2026-07-14 新建,但非本地模型42026-08-12见下方"2026 后继者判定"
duixcom/Duix-Reface❌ 404(Duix.Mobile README 里仍在推荐它,但已不可访问)——文档链接已失效
duixcom/Duix.mobile(旧名)301 → duixcom/Duix-Mobile——同为重命名

"2026 后继者是否取代 Duix.Heygem" 判定:

1.3 (c) 显存需求 + 16GB 单卡能否跑

1.4 (d) 端到端时延 / FPS

1.5 (e) 流式 / 可打断

1.6 (f) 中文支持 / 半身全身

1.7 (g) 坑:sm_120 / Blackwell / Docker 部署 / 许可

🔴 最致命:Blackwell 目前跑不起来(2026 年 9 月仍是 Open)

Issue #624,标题 "RTX 5070 (Blackwell) not supported — CUDA kernel error on Docker backend",opened on Sep 4, 2026,状态 Open,且页面零回复("Sign up for free to join this conversation")。原文逐字引用:

GPU: NVIDIA GeForce RTX 5070 Laptop GPU Compute Capability: 12.0 (Blackwell) Driver: 592.15 Error from duix-avatar-tts container: RuntimeError: CUDA error: no kernel image is available for execution on the device Root cause: The Docker images are compiled with a PyTorch version that doesn't include CUDA kernels for compute capability 12.0. PyTorch 2.7+ added Blackwell support. Request: Please rebuild and push updated Docker images (guiji2025/fish-speech-ziming, guiji2025/fun-asr, guiji2025/duix.avatar) compiled against PyTorch 2.7+ with sm_120 support.

链接:issue #624

→ 对 RTX 5060 Ti (sm_120) 的直接含义:默认三个 Docker 镜像的主线编译目标不含 sm_120,TTS 容器会直接抛 no kernel image is available。维护者截至 2026-09-22 未回应。

⚠️ 50 系有独立镜像,但只针对 5090,且未验证 sm_120/5060 Ti

部署形态:纯 Docker 打包,无源码级可调

三个服务(docker-compose-linux.yml):

服务镜像端口关键配置
duix-avatar-ttsguiji2025/fish-speech-ziming18180:8080runtime: nvidia,NVIDIA_VISIBLE_DEVICES=0
duix-avatar-asrguiji2025/fun-asr10095:10095runtime: nvidia,privileged: true
duix-avatar-gen-videoguiji2025/duix.avatar8383:8383runtime: nvidia,privileged: true,shm_size: '8g',PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512

坑点:

许可 / 商业授权 —— ⚠️ README 与 LICENSE 正文数字不一致,以 LICENSE 为准

我在本轮补抓了 LICENSE 正文(raw.githubusercontent.com/duixcom/Duix-Avatar/main/LICENSE,7120 字节),标题逐字为 DUIX.COM COMMUNITY LICENSE AGREEMENT(自定义社区许可,非标准 SPDX 许可)。

🔴 关键纠正:商用门槛是 1000 MAU,不是 README 说的 10 万用户。

LICENSE 里的其他附加义务(逐字摘要,易被忽略但都是硬性):

(对照:bytedance/LatentSync 的 LICENSE 正文我已抓取核实,为标准 Apache License Version 2.0,详见 3.6 节。)

同时 README 明说自部署版的定位是 "Lip Sync Effect: Usable effect"(可用级),而云端 API 才是 "Stunning and higher definition effect"、且 "Iteration Speed: Slow updates, bug fixes depend on the community" —— 官方自认开源版迭代慢。


2. Rudrabha/Wav2Lip

2.1 2026 状态:事实性停更(effectively abandoned)

项目数值证据
Stars / Forks13219 / 2849ungh.cc/repos/Rudrabha/Wav2Lip
创建2020-08-07同上
最后 push2025-06-22同上。→ 2026 年零更新
最新 Issue 活动搜索页里最新的 Issue 是 #613 "Extracting raw audio... and stuck. Load not working",opened on Jan 3, 2024;#584 opened on Nov 10, 2023issues 搜索

停更的硬证据(README 已被改成商业导流页):master 分支 README 开头第一段不再是论文说明,而是:

Commercial Version

Create your first lipsync generation in minutes. Please note, the commercial version is of a much higher quality than the old open source model! Create your API key from the Dashboard.

其后大段内容是 Sync.so 的 pip install syncsdk 调用示例,论文原文被挤到 README 尾部。(README) → 结合 2025-06-22 那次 push 的性质(把 README 换成商业 API 导流,而非代码更新),以及 Issue 追踪器自 2024 年初起基本无维护者回应,可以判定:Wav2Lip 在 2026 年已是"存档态"项目,作者精力已转移到商业产品 sync.so。

社区推荐的后继者:

2.2 (c) 显存需求 + 16GB

2.3 (d) 端到端时延 / FPS

未查到官方或可引用的实测数字。 唯一相关线索是 Issue #584(标题 "How much time do you need to lip sync a 10 sec or 1 minute video?",opened on Nov 10, 2023,至今 Open 且无人给出有效回答),提问者原文:

I have been trying the last days with both wav2lip HD (not in auto) and retalker, and found that both are slow and very GPU consuming. I would like to know everyone of you HOW MUCH GPU do you use (what card) and HOW MUCH time does it take for you to do it? Please contribute. Because I am about to drop this technology and give up on it, maybe others peoples experiences will give me hope.

链接:issue #584 → 该 Issue 本身没有给出数字,它只能证明"用户普遍抱怨慢、且社区无人应答"。

2.4 (e) 流式 / 可打断

不支持。 架构是整段批处理:python inference.py --checkpoint_path <ckpt> --face <video.mp4> --audio <an-audio-source>,结果一次性写盘为 results/result_voice.mp4。无流式接口、无 chunk 级回调、不可中途打断。(README)

2.5 (f) 中文支持 / 半身全身

2.6 (g) 坑:sm_120 / 依赖 / 维持性

librosa==0.7.0
numpy==1.17.1
opencv-contrib-python>=4.2.0.34
opencv-python==4.1.0.25
torch==1.1.0
torchvision==0.3.0
tqdm==4.45.0
numba==0.48

PyPI 硬证据(curl https://pypi.org/pypi/torch/1.1.0/json):torch==1.1.0 全仓仅 9 个 wheel,上传时间 2019-04-30,Python 标签只有 cp27 / cp35 / cp36 / cp37——没有任何 cp38+ 轮子,文件名为 torch-1.1.0-cp37-cp37m-manylinux1_x86_64.whl(manylinux1 是 2010 年 ABI 基线)。同批老矩阵:torchvision==0.3.0(2019-05-22)、numpy==1.17.1(2019-08-27)、numba==0.48(2020-06-30)。 三层含义:① 装不上 —— Ubuntu 24.04 是 Python 3.12,torch 1.1.0 无对应 wheel,pip install -r requirements.txt 直接失败;② 跑不了 —— 1.1.0 的 CUDA 轮子属 CUDA 9/10 时代,arch 上限远低于 sm_120,必报 no kernel image available;③ 生态全断 —— numba 0.48 / numpy 1.17.1 / librosa 0.7.0 在 Python 3.12 上均无可用轮子。 → 判定:该 requirements.txt 在 2026 年 Ubuntu + sm_120 上属于"不可复现文档",只能当算法参考,必须用现代 torch 重写推理脚本(社区各 fork / ComfyUI 节点本质都是这么做的)。这也解释了为何 Issue 追踪器自 2024 初就没人回答环境问题。


3. bytedance/LatentSync

3.1 2026 状态与版本

项目数值证据
Stars / Forks6090 / 981ungh.cc/repos/bytedance/LatentSync;GitHub Issues 页 aria-label 亦为 "6090 users starred this repository"
创建 / 最后 push2024-12-11 / 2025-06-20同上 → 2026 年零代码更新(截至 2026-09-22 已停更约 15 个月)
GitHub Releases空({"releases":[]})同上。权重走 HuggingFace,不打 GitHub Release
最新版本1.6(2025/06/11)README Updates:"2025/06/11: We released LatentSync 1.6, which is trained on 512×512 resolution videos to mitigate the blurriness problem.";HF 仓 ByteDance/LatentSync-1.6

"2026 后继版是否改变要求" 的答案:没有 2026 新版;而 1.6 相对 1.5 把显存要求改坏了(8GB → 18GB)。

3.2 (c) 显存需求 —— 16GB 单卡的判决书就在这里

【作者自述,README 原文,最权威】

Minimum VRAM for inference:

  • 8 GB with LatentSync 1.5
  • 18 GB with LatentSync 1.6

链接:LatentSync README

→ 16GB 单卡 < 官方 18GB 门槛,官方口径下 1.6 不可行。 要留在 16GB 内,官方口径只允许退回 1.5(8GB)。

训练档(同一 README,供参考):stage1.yaml 23GB;stage2.yaml 30GB;stage2_efficient.yaml 20GB(作者注:"suitable for users with consumer-grade GPUs, such as the RTX 3090");stage1_512.yaml 30GB;stage2_512.yaml 55GB。(README)

fp16/权重体积(第三方):ComfyUI 封装仓 README 列出 latentsync_unet.pt (~5GB)、stable_syncnet.pt (~1.6GB),并称 "Reduced VRAM Requirements: Optimized to run on 20GB VRAM (RTX 3090 compatible)"。(ShmuelRonen/ComfyUI-LatentSyncWrapper)⚠️ 该仓为第三方封装,且"20GB / RTX 3090"表述自相矛盾(3090 是 24GB),仅作量级参考。

🟢 决定性第三方实测:RTX 5070 Ti 16GB 真的跑起来了(2026 年 7 月) Issue #365 "TensorRT-RTX optimized inference path and reduced-step speed benchmarks",opened on Jul 3, 2026,作者 Petrus Vermaak(自述 "This work was directed, tested, and validated by Petrus Vermaak... OpenAI Codex served as the implementation and engineering agent")。测试系统:RTX 5070 Ti 16 GB test system —— 同为我们关心的显存档位与 Blackwell 世代。基准夹具原文:

Fixture: official demo video/audio, 9.68s output, 242 frames, 512 face processing, 20 steps, guidance 1.5.

BackendWall secondsSec/output secondChange vs exact20Peak total VRAM
exact20 PyTorch eager397.51241.065baseline15807 MB
optimized20 TensorRT-RTX exact268.16027.70232.54% faster12551 MB

(逐字引用表格数字;链接:issue #365)

→ 【第三方实测】16GB 单卡在 512×512 / 20 步下可以跑完 LatentSync 1.6:裸 PyTorch 峰值 15807 MB(已占满 16GB 的 96%,极限)、TensorRT-RTX 优化后 12551 MB(留有余量)。官方说的 18GB 是保守门槛,实测 16GB 能压进去,但 PyTorch 裸跑基本贴着天花板。

用户侧旁证(都指向 16GB 不够):

3.3 (d) 端到端时延 / FPS

3.4 (e) 流式 / 可打断

不支持,且被 Issue 明确暴露。 架构是整段扩散去噪(20 步),没有 chunk 级流水线;Issue #326 标题即为 "How to run inference on longer video?"(issue #326)→ 长视频要自己切段处理。无流式输出、无打断机制。

3.5 (f) 中文支持 / 半身全身

3.6 (g) 坑:torch/CUDA / xformers / flash-attn / 512×512 显存

🔴 最硬的坑:官方依赖锁死在 cu121,不含 sm_120 内核。 requirements.txt 逐字(raw):

torch==2.5.1
torchvision==0.20.1
--extra-index-url https://download.pytorch.org/whl/cu121
diffusers==0.32.2
transformers==4.48.0
decord==0.6.0
accelerate==0.26.1
einops==0.7.0
omegaconf==2.3.0
opencv-python==4.9.0.80
mediapipe==0.10.11
python_speech_features==0.6
librosa==0.10.1
scenedetect==0.6.1
ffmpeg-python==0.2.0
imageio==2.31.1
imageio-ffmpeg==0.5.1
lpips==0.1.4
face-alignment==1.4.1
gradio==5.24.0
huggingface-hub==0.30.2
numpy==1.26.4
kornia==0.8.0
insightface==0.7.3
onnxruntime-gpu==1.21.0
DeepCache==0.1.1

→ torch==2.5.1 + cu121 官方 wheel 不含 sm_120(Blackwell)内核。在 RTX 5060 Ti 上照原样装,会撞上与 Duix #624 同型的 no kernel image is available for execution on the device。必须自行升级到 torch ≥ 2.7 + cu128/cu13x,而任何 torch 升级都可能打破下面这批老旧 pin。这是本项目在 sm_120 上最大的工作量来源。

✅ 好消息:不需要 xformers,也不需要 flash-attn。 我逐条核对了上面的 requirements:没有 xformers,也没有 flash-attn。→ 不存在 flash-attn/triton 需要现场编译 sm_120 内核的问题,这一点比很多扩散类 talking-head 项目(如 MuseTalk / 部分 ComfyUI 节点)友好得多。注意力走 diffusers 默认路径。 (若走 TensorRT-RTX 路线,Issue #365 已证明在 16GB Blackwell 上可行且比 PyTorch 省 3256MB 峰值显存、快 32.54%,是本项目在 16GB 卡上最值得走的优化路径。)

其他版本坑(逐条):

3.7 LatentSync 面向 RTX 5060 Ti 16GB 的一句话判决

能跑,但要打破官方依赖并接受"离线批处理"定位:官方 18GB 门槛把 1.6 判为不合规;实测 16GB(5070 Ti)在 512×512/20 步下峰值 15807MB 可通过,TensorRT-RTX 降到 12551MB。代价是速度仅约实时的 1/28~1/41,且必须自己把 torch 从 cu121 迁到 cu128+ 才能吃 sm_120。 它不是"实时数字人",是"离线 lip-sync 渲染器"。


4. 三仓横向对照(面向 RTX 5060 Ti 16GB / sm_120)

维度Duix.Avatar(原 Duix.Heygem)Wav2LipLatentSync
2026 状态改名后仍活跃(2026-04 有提交)停更(2025-06-22 最后一次 push,README 已商业导流)停更 15 个月(2025-06-20)
Stars15547132196090
官方推理显存未给数字(整机推荐 RTX 4070,16GB+ 为推荐档)未给8GB(1.5) / 18GB(1.6)
16GB 单卡⚠️ 怕是不行:三容器常驻,无 16GB 实测✅ 几无悬念可跑(推断)⚠️ 官方不合规,实测 15807MB 压线可跑
成片速度未查到 FPS(官方定义非实时)未查到(Issue 普遍抱怨慢)27.7~41.1 秒 / 1 秒成片(5070Ti 16GB 实测)
流式/打断❌ 提交+轮询❌ 整段批处理❌ 整段扩散
中文✅ 8 语言含中文⚠️ 语言无关,但只训练英文(LRS2)✅ 明确优化中文视频
半身/全身【推断】继承源视频(face2face)只改嘴部,身体不动❌ 仅面部
sm_120 现状🔴 坏(Issue #624 Open 零回复,主线镜像无 CC12.0 内核;仅 5090 专用镜像存在,未验证 5060 Ti)🔴 pin 即不可用(torch==1.1.0 为 2019 年 CUDA 9/10 时代,且无 cp38+ 轮子;须整套重写)🟡 需自行把 cu121→cu128+(无 xformers/flash-attn,省一大坑)
商用许可⚠️ 自定义许可(非开源):DUIX.COM COMMUNITY LICENSE,MAU > 1000 即须申请商业授权(README 误写"10 万",差 100 倍)❌ 严格禁止商用(LRS2 数据集限制)✅ Apache License 2.0(LICENSE 正文已核实)
部署形态Docker 三容器 + 客户端 App裸 Python 脚本裸 Python 脚本

5. 依赖 pin 对 sm_120 / Blackwell 的含义(本轮重点补充)

5.1 事实基线:要原生支持 sm_120 需要什么

【官方,PyTorch 官方博客】PyTorch 2.7 发布说明逐字(pytorch.org/blog/pytorch-2.7,180248 字节,我已抓取正文):

→ 基线结论:要原生吃 sm_120,torch 必须 ≥ 2.7 且用 cu128(或更新)轮子;cu121 及更早的官方轮子不含 sm_120 内核。

【第三方实测,逐字警告文本 —— 索引标题级证据】 Comfy-Org/ComfyUI Issue #7127 的页面标题逐字包含运行期警告:

UserWarning: NVIDIA GeForce RTX 5070 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation. The current PyTorch install supports CUDA capabilities sm_50 sm_60 sm_61 sm_70 sm_75 sm_80 sm_86 sm_90.

→ 这就是 cu121 轮子的实际 arch 列表:最高到 sm_90,没有 sm_120。触发设备是 RTX 5070 Ti(同为 sm_120),与我们的 5060 Ti 同代同 cap。 ⚠️ 诚实标注:本轮 github.com 的 Issue 正文页多次抓取超时(网络问题),上述引文来自搜索索引返回的页面标题原文,我未读取该 Issue 正文;pytorch/pytorch #166794 同理,其标题逐字为 "[Bug] RTX 5070 Ti (sm_120) not recognized by PyTorch 2.5.1+cu121"。两者作为"同型 pin 已被报告失败"的证据成立,但均为标题级、非正文级证据。 【第三方实测,同型】HybridRobotics/Berkeley-Humanoid-Lite Issue #49 标题逐字:"RTX 5090 (SM_120) fails with default torch/cu121: no kernel image is available"。

5.2 三条路线的 pin 逐条判定

仓库pin(逐字)sm_120 可用性必须做什么
LatentSynctorch==2.5.1 / torchvision==0.20.1 / --extra-index-url .../cu121 / onnxruntime-gpu==1.21.0🔴 不可用:2.5.1 早于 2.7,cu121 arch 表止于 sm_90升 torch ≥2.7 + cu128,并连带处理 5.4 的脆点
Wav2Liptorch==1.1.0 / torchvision==0.3.0 / numba==0.48 / numpy==1.17.1 / librosa==0.7.0🔴 绝对不可用:2019 年 CUDA 9/10 产物,且无 cp38+ 轮子pin 形同废弃,须整体现代化重写
Duix.Avatar源码无 pin——pin 在 Docker 镜像内部🔴 已实测失败:Issue #624 证实镜像不含 CC 12.0 内核用户无法自行解决,只能等官方重建镜像或自建镜像

特别注意 LatentSync 与 #166794 的 pin 是同型:torch 2.5.1 + cu121 正是被反复报告在 RTX 5070 Ti / 5090(sm_120) 上失败的那一组。这不是理论风险,而是已在同代硬件上被复现的失败组合。

5.3 Wav2Lip 的 torch==1.1.0 到底有多致命(本轮新查,PyPI 硬证据)

【第三方可核实,PyPI JSON API】torch==1.1.0 的全部 9 个 wheel(https://pypi.org/pypi/torch/1.1.0/json):

三层含义:

  1. 装不上:Ubuntu 24.04 默认 Python 3.12,torch 1.1.0 无对应 wheel,pip install -r requirements.txt 必然失败;若强行从源码编译,2019 年的 CUDA 代码在 CUDA 13.2 的 nvcc 下几乎必然编译失败。
  2. 跑不了:即便人工凑出环境,1.1.0 的 CUDA 轮子属 CUDA 9.0/10.0 时代,arch 覆盖远低于 sm_120,必报 no kernel image is available for execution on the device(与 Duix #624 同型错误)。
  3. 生态全断:numba==0.48(依赖 llvmlite 0.31)、numpy==1.17.1、librosa==0.7.0 在 Python 3.12 上均无可用轮子。

→ 最终判定:Wav2Lip 的 requirements.txt 在 2026 年的 Ubuntu + sm_120 上属于"不可复现文档"。它无法按官方方式部署,只能把 wav2lip.pth 权重 + 推理逻辑搬到现代 torch 上重写(社区各种 fork 与 ComfyUI 节点本质都是这么做的)。这与 README 自述的 Python 3.6 前提互相印证,也解释了为何其 Issue 追踪器自 2024 年初起就无人处理环境类问题。

5.4 LatentSync 换 torch 时需连带处理的 5 个脆点

唯一的硬门槛是 torch 2.5.1 → ≥2.7(+cu128),但换版本号会牵动一串 2025 年上半年的旧 pin:

  1. cuDNN 大版本:不冲突(好消息)。ORT 文档逐字:"PyTorch 2.3 uses cuDNN 8.x, while PyTorch 2.4 or later uses cuDNN 9.x" → 2.5.1 与 2.7 同为 cuDNN 9.x,这一项不需要额外折腾。
  2. decord==0.6.0:视频解码轮子,长期只发布到 cp310/cp311 附近,新 Python 上常无轮子 → 这类项目的经典断点。
  3. insightface==0.7.3:需现场 C++ 编译(人脸检测依赖),是环境搭建的常见失败点。
  4. DeepCache==0.1.1 + diffusers==0.32.2:DeepCache 是推理加速补丁,与新版 diffusers 的组合从未被作者验证(作者 2025-06-20 后停更)。
  5. onnxruntime-gpu==1.21.0:与你的 CUDA 13.2 宿主最容易打架,单独见 5.5。

✅ 一个重要的好消息:LatentSync 的 requirements.txt 没有 xformers,也没有 flash-attn → 不需要为 sm_120 现场编译 flash-attn / triton 内核。中文资料里常见的那类报错(如 cnblogs 记录的 "nvcc fatal : Unsupported gpu architecture 'compute_120'")在 LatentSync 的原生路径上不会出现,只在你想额外加装 flash-attn 时才会撞上。若走 Triton 路线,PyTorch 2.7 已内置 Triton 3.3,官方称其 "adds support for the Blackwell architecture with torch.compile compatibility"。

5.5 onnxruntime-gpu==1.21.0 与 CUDA 13.2 宿主的冲突(本轮新查)

【官方,ONNX Runtime CUDA EP 文档逐字】(onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider,正文已抓取):

→ onnxruntime-gpu==1.21.0 是 CUDA 12.8 + cuDNN 9.x 构建,而你的宿主是 CUDA 13.2(跨大版本)。按 NVIDIA Minor Version Compatibility,12.8 构建只保证与 12.x 兼容。 → 【推断,风险项】两条路:① 额外提供 CUDA 12 运行库(如 pip 的 nvidia-cuda-runtime-cu12 / nvidia-cudnn-cu12),否则可能缺 libcudart.so.12;② 把 ORT 升到 ≥1.27(cu13.0 构建)以匹配宿主 —— 但后者已脱离 LatentSync 的 pin,必须重测。 → Blackwell 专门适配:搜索索引显示 microsoft/onnxruntime 有一个 PR #23928 "Extend CMAKE_CUDA_FLAGS with all Blackwell compute capacity"(标题逐字),说明 ORT 曾专门为 Blackwell 补齐 CUDA arch 编译标志。⚠️ 该 PR 进入哪个发布版本我未能核实(PR 页本轮网络不可达),标注为未核实;实务上 cu12.8 构建通常可借 PTX JIT 在 sm_120 上跑通,但需实测确认。

5.6 汇总:三条路线在 RTX 5060 Ti 16GB 上的 sm_120 工作量

仓库16GB 上能否跑改动量主要拦路虎
LatentSync✅ 实测可行(峰值 15807MB 压线 / TensorRT 后 12551MB)中:换 torch≥2.7+cu128,处理 decord / ORT / DeepCache无 flash-attn 坑是最大优势;ORT 与 CUDA 13.2 需调和
Duix.Avatar❌ 目前不可行(#624 已实测同型失败)无法自行解决(内核编译在官方镜像里)官方 #624 零响应;仅 5090 专用镜像,未验证 5060 Ti
Wav2Lip⚠️ 理论上限低但无法按官方方式部署大:2019 依赖整套作废,须重写推理脚本torch==1.1.0 无 cp38+ 轮子;且严格禁止商用

一句话:三个项目的 pin 无一为 sm_120 准备。LatentSync 只差一个 torch 升级(且没有 flash-attn 这个天坑),是三者中唯一在 16GB Blackwell 上被实测跑通的;Duix 卡在官方镜像、用户无力自救;Wav2Lip 的依赖本身就是 2019 年化石。


来源清单

Duix 系

  1. https://github.com/duixcom/Duix-Avatar — 主仓(原 Duix.Heygem),README 硬件/语言/许可/API 口径
  2. https://github.com/duixcom/Duix-Avatar/releases — v1.0.5 / v1.0.6 改名说明
  3. https://github.com/duixcom/Duix-Avatar/issues/624 — RTX 5070 Blackwell no kernel image is available(2026-09-04 开,Open 零回复)
  4. https://github.com/duixcom/Duix-Avatar/issues/325 — "6G VRAM can't run with lite version"
  5. https://github.com/duixcom/Duix-Avatar/issues/251 — 4×24G 显卡推理慢
  6. https://raw.githubusercontent.com/duixcom/Duix-Avatar/main/deploy/docker-compose-linux.yml — 三服务定义
  7. https://raw.githubusercontent.com/duixcom/Duix-Avatar/main/deploy/docker-compose-5090.yml — 50 系专用镜像
  8. https://deepwiki.com/duixcom/Duix-Avatar/6.3-hardware-and-performance-tuning — 硬件/显存表(AI 生成 Wiki,含 5090 = CC 9.0 错误,仅供参照)
  9. https://github.com/duixcom/Duix-Mobile — 移动端实时 SDK(120ms / <1.5s / 流式 barge-in)
  10. https://github.com/duixcom/duix-skills — 2026-07-14 新建的云端 API skill(非本地后继者)
  11. https://ungh.cc/repos/duixcom/Duix-Avatar — stars/日期(2026-09-22 读取)
  12. https://ungh.cc/repos/duixcom/Duix-Mobile 、 https://ungh.cc/repos/duixcom/duix-skills — 同上

Wav2Lip

  1. https://github.com/Rudrabha/Wav2Lip — README(商业导流、语言、许可、Python 3.6)
  2. https://github.com/Rudrabha/Wav2Lip/issues/584 — 10s/1min 耗时提问(2023-11-10,Open 无答)
  3. https://github.com/Rudrabha/Wav2Lip/issues?q=is%3Aissue+VRAM — Issue 追踪器活动停滞证据
  4. https://ungh.cc/repos/Rudrabha/Wav2Lip — stars / 最后 push 2025-06-22

LatentSync

  1. https://github.com/bytedance/LatentSync — README(1.5=8GB / 1.6=18GB、训练档、中文优化、面部仿射管线、版本更新日志)
  2. https://raw.githubusercontent.com/bytedance/LatentSync/main/requirements.txt — torch==2.5.1 + cu121(无 xformers/flash-attn)
  3. https://github.com/bytedance/LatentSync/issues/365 — RTX 5070 Ti 16GB 实测:397.512s/15807MB、268.160s/12551MB(2026-07-03)
  4. https://github.com/bytedance/LatentSync/issues/335 — "因為顯存不夠,只有16GB"
  5. https://github.com/bytedance/LatentSync/issues/314 — "8G显存运行不了1.5"
  6. https://github.com/bytedance/LatentSync/issues/284 — "v1.6 是否可以在8G vRam 运行?"
  7. https://github.com/bytedance/LatentSync/issues/137 — 4.17it/s、GPU 利用率 1%-3%、6GB/16GB
  8. https://github.com/bytedance/LatentSync/issues/270 — "No available kernel. Aborting execution."
  9. https://github.com/bytedance/LatentSync/issues/326 — 长视频推理
  10. https://huggingface.co/ByteDance/LatentSync-1.6 — 1.6 权重(最新版本)
  11. https://github.com/ShmuelRonen/ComfyUI-LatentSyncWrapper — 第三方封装(~5GB unet / ~1.6GB syncnet / "20GB VRAM" 表述)
  12. https://ungh.cc/repos/bytedance/LatentSync — stars / 最后 push 2025-06-20

依赖 pin / sm_120 / 许可 专项(本轮补查)

  1. https://raw.githubusercontent.com/Rudrabha/Wav2Lip/master/requirements.txt — 逐字 torch==1.1.0 / torchvision==0.3.0 / numba==0.48 / numpy==1.17.1 / librosa==0.7.0
  2. https://pypi.org/pypi/torch/1.1.0/json — torch 1.1.0 全部 9 个 wheel,上传 2019-04-30,仅 cp27/cp35/cp36/cp37
  3. https://pypi.org/pypi/torchvision/0.3.0/json 、 https://pypi.org/pypi/numba/0.48/json 、 https://pypi.org/pypi/numpy/1.17.1/json — 同批 2019–2020 上传时间
  4. https://raw.githubusercontent.com/duixcom/Duix-Avatar/main/LICENSE — DUIX.COM COMMUNITY LICENSE AGREEMENT 正文;第 2 条 MAU > 1 thousand 门槛;第 1.b.i / 5.c / 5.b 条附加义务
  5. https://raw.githubusercontent.com/bytedance/LatentSync/main/LICENSE — Apache License Version 2.0 正文(已核实)
  6. https://pytorch.org/blog/pytorch-2.7/ — "PyTorch 2.7 introduces support for NVIDIA's new Blackwell GPU architecture and ships pre-built wheels for CUDA 12.8";"[Prototype] NVIDIA Blackwell Architecture Support";Triton 3.3
  7. https://github.com/Comfy-Org/ComfyUI/issues/7127 — UserWarning 逐字含 cu121 的 arch 列表 sm_50 ... sm_90(无 sm_120)(索引标题级证据)
  8. https://github.com/pytorch/pytorch/issues/166794 — "[Bug] RTX 5070 Ti (sm_120) not recognized by PyTorch 2.5.1+cu121"(索引标题级证据,与 LatentSync pin 同型)
  9. https://github.com/HybridRobotics/Berkeley-Humanoid-Lite/issues/49 — "RTX 5090 (SM_120) fails with default torch/cu121: no kernel image is available"(索引标题级证据)
  10. https://onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider — ORT 1.21.x–1.26.x = CUDA 12.8 + cuDNN 9.x;1.27 起默认 CUDA 13.0;12.8 构建仅保证与 CUDA 12.x 兼容
  11. https://github.com/microsoft/onnxruntime/pull/23928 — "Extend CMAKE_CUDA_FLAGS with all Blackwell compute capacity"(索引标题级证据;所入版本未核实)

未查到 / 未能核实的项(诚实标注)

下载此文件