【作者自述】=仓库 README / Release 官方口径;【第三方实测】=Issue / 博客 / Wiki 中他人给出的数字;【推断】=基于文件与文档的合理推导,非直接证据;未查到=未能验证,不编造。duixcom/Duix.Heygem 这个仓库在 2026 年已经不存在了 —— 它被官方重命名为 duixcom/Duix-Avatar,GitHub 对旧地址做 301 跳转。所以"Duix.Heygem 2026 状态"的正确答案是:主体已改名,改名后的 Duix-Avatar 是唯一活跃主线。CUDA error: no kernel image is available for execution on the device,根因指向 Docker 镜像里的 PyTorch 不含 CC 12.0 内核,该 Issue 至今 Open 且零回复。| 项目 | 结果 | 证据 |
|---|---|---|
duixcom/Duix.Heygem | 301 跳转到 duixcom/Duix-Avatar | curl -sL -w '%{url_effective}' → 200 -> https://github.com/duixcom/Duix-Avatar;ungh.cc 对 duixcom/Duix.Heygem 查询返回 "id":907627874,"name":"Duix-Avatar"(同一个 repo id,证明是重命名而非新仓) |
| Stars / Forks | 15547 / 2642 | ungh.cc/repos/duixcom/Duix-Avatar,2026-09-22 读取 |
| 创建 / 最后 push | 创建 2024-12-24T03:00:26Z;最后 push 2026-04-21T07:06:36Z | 同上。→ 2026 年有提交(4 月),但到 9 月已约 5 个月无代码更新 |
| 改名证据(Release) | v1.0.5(2025-08-15)发布说明:"The original project name "HeyGem" has now been officially changed to "Duix.Avatar".";v1.0.6(2025-09-28):"It's just a name change. If the version you're using has no issues, you can continue to use it." | releases |
duixcom org 下各仓 2026 状态(逐个核查,含"是否存在")说明:任务要求的
curl https://api.github.com/orgs/duixcom/repos因本机 IP 限流(core remaining=0,reset 13:52)无法执行;改用 ungh.cc 镜像逐个精确查询 + 落地页跳转探测,结论同样可证伪。
| 仓库 | 状态(2026-09-22) | Stars | 最后 push | 说明 |
|---|---|---|---|---|
duixcom/Duix-Avatar(原 Duix.Heygem) | ✅ 存在、2026 有更新 | 15547 | 2026-04-21 | 本组主目标的正身 |
duixcom/Duix.Mobile(原 Duix.mobile) | ✅ 存在、2026 活跃 | 8252 | 2026-08-05 | 移动端/嵌入式实时交互 SDK,非本地 GPU 服务。仓描述:"on-premise deployment and <1.5 s latency" |
duixcom/Duix.Heygem.Android | ❌ 404,不存在 | — | — | 该仓从未存在或已删除;raw 的 main/master 均 404 |
duixcom/duix-skills | ⚠️ 2026-07-14 新建,但非本地模型 | 4 | 2026-08-12 | 见下方"2026 后继者判定" |
duixcom/Duix-Reface | ❌ 404(Duix.Mobile README 里仍在推荐它,但已不可访问) | — | — | 文档链接已失效 |
duixcom/Duix.mobile(旧名) | 301 → duixcom/Duix-Mobile | — | — | 同为重命名 |
"2026 后继者是否取代 Duix.Heygem" 判定:
duixcom/duix-skills(2026-07-14 新建)不构成本地部署的后继者:它是一套给 AI agent 用的云端 API skill,需要 DUIX_APP_ID/DUIX_APP_KEY/DUIX_API_KEY 凭据,走订阅套餐 + 预付费积分计费,产出是 conversation_url 网页链接或云端生成的 MP4;其 README 明确写 "Prepaid credits (consumed by task video duration)"、"[Production time] About 20 minutes to 2 hours"。→ 它把 Duix 的 lip-sync 能力 SaaS 化了,与"本地 16GB 显卡部署"是两条路。(证据:duix-skills README)CPU: 13th Gen Intel Core i5-13400F / Memory: 32GB / Graphics Card: RTX 4070,硬盘 C 盘 >100GB、D 盘 >30GB;服务器部署要下载 约 70GB 流量。(README)GPU | NVIDIA RTX 4070 (8GB VRAM) | NVIDIA RTX 4090 / RTX 5090 (16GB+ VRAM)(Minimum → Recommended),并给出调优表:8GB → max_split_size_mb 256/512、12-16GB → 512 (default)、24GB+ → 1024。 ⚠️ 该页有硬错误需警惕:它把 RTX 5090 的 "CUDA Compute" 标为 9.0(5090 实为 sm_120),说明此页为 AI 自动生成、未经人工校对;其数字仅可作参考。链接:deepwiki.com/duixcom/Duix-Avatar/6.3-hardware-and-performance-tuningguiji2025/duix.avatar),仓库不发布独立 .pt/.safetensors,故无 fp16 权重体积数字可引用。Docker Hub 我尝试拉取 tag 体积,多次超时未成功(hub.docker.com 在本机网络下不可达)。Duix.Mobile 的时延数字不能用于 Duix.Avatar —— Mobile README 写 "AI avatar response latency under 120ms (tested on Snapdragon® 8 Gen 2 SoC)",那是手机 SoC 上的端侧渲染,与 GPU 服务的视频合成是两套东西。(Duix-Mobile README)http://127.0.0.1:8383/easy/submit,进度查询 http://127.0.0.1:8383/easy/query?code=${taskCode}(GET 轮询);TTS 接口参数里甚至有 "streaming": false, // Fixed parameter 这一固定值。→ 非流式、不可打断。(README)~/duix_avatar_data/face2face:/code/data,即在源视频上重绘人脸区域。据此推断输出构图继承源视频,源视频是半身/全身就得到半身/全身;但这属于推断,非文档明示,我未查到官方对"半身/全身"的明确能力声明。证据:deploy/docker-compose-linux.ymlIssue #624,标题 "RTX 5070 (Blackwell) not supported — CUDA kernel error on Docker backend",opened on Sep 4, 2026,状态 Open,且页面零回复("Sign up for free to join this conversation")。原文逐字引用:
GPU: NVIDIA GeForce RTX 5070 Laptop GPU Compute Capability: 12.0 (Blackwell) Driver: 592.15 Error from duix-avatar-tts container:
RuntimeError: CUDA error: no kernel image is available for execution on the deviceRoot cause: The Docker images are compiled with a PyTorch version that doesn't include CUDA kernels for compute capability 12.0. PyTorch 2.7+ added Blackwell support. Request: Please rebuild and push updated Docker images (guiji2025/fish-speech-ziming,guiji2025/fun-asr,guiji2025/duix.avatar) compiled against PyTorch 2.7+ with sm_120 support.
链接:issue #624
→ 对 RTX 5060 Ti (sm_120) 的直接含义:默认三个 Docker 镜像的主线编译目标不含 sm_120,TTS 容器会直接抛 no kernel image is available。维护者截至 2026-09-22 未回应。
deploy/docker-compose-5090.yml(HTTP 200,1064 字节),镜像换成了 guiji2025/duix.avatar-5090 与 guiji2025/fish-speech-5090,仍保留 PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512 与 shm_size: '8g'。 → 【推断】50 系走单独镜像这条路径"理论上"是给 Blackwell 的,但官方只声明在 5090 上验证过;在 5060 Ti 上是否可用未查到任何证据,且 #624 的存在说明普通镜像确实没有 sm_120 内核。这是 16GB 选型里最大的未知风险点。 (证据:docker-compose-5090.yml)三个服务(docker-compose-linux.yml):
| 服务 | 镜像 | 端口 | 关键配置 |
|---|---|---|---|
| duix-avatar-tts | guiji2025/fish-speech-ziming | 18180:8080 | runtime: nvidia,NVIDIA_VISIBLE_DEVICES=0 |
| duix-avatar-asr | guiji2025/fun-asr | 10095:10095 | runtime: nvidia,privileged: true |
| duix-avatar-gen-video | guiji2025/duix.avatar | 8383:8383 | runtime: nvidia,privileged: true,shm_size: '8g',PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512 |
坑点:
runtime: nvidia,其中两个 privileged: true → 与宿主驱动强耦合;README 明确 NVIDIA 显卡与驱动是硬性前提,且 FAQ 原文 "All computing power of this project is local. The three services won't start without an NVIDIA graphics card or proper drivers."sudo nvidia-ctk runtime configure --runtime=docker(README 给了完整步骤);没有 CPU 回退路径(DeepWiki:"The system has no CPU-only fallback; an NVIDIA GPU is mandatory for all deployment modes.")shm_size: '8g' 是硬需求:DeepWiki 把 "Bus Error (Core Dumped)" 归因于 "Insufficient shared memory → Increase shm_size to 12g or 16g";1080p/60s 就建议 8GB,4K 建议 12GB+。我在本轮补抓了 LICENSE 正文(raw.githubusercontent.com/duixcom/Duix-Avatar/main/LICENSE,7120 字节),标题逐字为 DUIX.COM COMMUNITY LICENSE AGREEMENT(自定义社区许可,非标准 SPDX 许可)。
🔴 关键纠正:商用门槛是 1000 MAU,不是 README 说的 10 万用户。
LICENSE 里的其他附加义务(逐字摘要,易被忽略但都是硬性):
Notice 文本文件,含 "DUIX.COM is licensed under the DUIX.COM Community License, Copyright © DUIX.COM Platforms, Inc. All Rights Reserved."(对照:bytedance/LatentSync 的 LICENSE 正文我已抓取核实,为标准 Apache License Version 2.0,详见 3.6 节。)
同时 README 明说自部署版的定位是 "Lip Sync Effect: Usable effect"(可用级),而云端 API 才是 "Stunning and higher definition effect"、且 "Iteration Speed: Slow updates, bug fixes depend on the community" —— 官方自认开源版迭代慢。
| 项目 | 数值 | 证据 |
|---|---|---|
| Stars / Forks | 13219 / 2849 | ungh.cc/repos/Rudrabha/Wav2Lip |
| 创建 | 2020-08-07 | 同上 |
| 最后 push | 2025-06-22 | 同上。→ 2026 年零更新 |
| 最新 Issue 活动 | 搜索页里最新的 Issue 是 #613 "Extracting raw audio... and stuck. Load not working",opened on Jan 3, 2024;#584 opened on Nov 10, 2023 | issues 搜索 |
停更的硬证据(README 已被改成商业导流页):master 分支 README 开头第一段不再是论文说明,而是:
Commercial Version
Create your first lipsync generation in minutes. Please note, the commercial version is of a much higher quality than the old open source model! Create your API key from the Dashboard.
其后大段内容是 Sync.so 的 pip install syncsdk 调用示例,论文原文被挤到 README 尾部。(README) → 结合 2025-06-22 那次 push 的性质(把 README 换成商业 API 导流,而非代码更新),以及 Issue 追踪器自 2024 年初起基本无维护者回应,可以判定:Wav2Lip 在 2026 年已是"存档态"项目,作者精力已转移到商业产品 sync.so。
社区推荐的后继者:
rudrabha@synclabs.so / prajwal@synclabs.so 与 https://synclabs.so/)。Python 3.6、ffmpeg。(README)Wav2Lip 与 Wav2Lip + GAN 两个 checkpoint),仓库内不发布文件;我尝试检索 CSDN/Reddit 的实测数字,未获得可引用的具体数值,故不编造。可确认的只有量级判断:未查到官方或可引用的实测数字。 唯一相关线索是 Issue #584(标题 "How much time do you need to lip sync a 10 sec or 1 minute video?",opened on Nov 10, 2023,至今 Open 且无人给出有效回答),提问者原文:
I have been trying the last days with both wav2lip HD (not in auto) and retalker, and found that both are slow and very GPU consuming. I would like to know everyone of you HOW MUCH GPU do you use (what card) and HOW MUCH time does it take for you to do it? Please contribute. Because I am about to drop this technology and give up on it, maybe others peoples experiences will give me hope.
链接:issue #584 → 该 Issue 本身没有给出数字,它只能证明"用户普遍抱怨慢、且社区无人应答"。
不支持。 架构是整段批处理:python inference.py --checkpoint_path <ckpt> --face <video.mp4> --audio <an-audio-source>,结果一次性写盘为 results/result_voice.mp4。无流式接口、无 chunk 级回调、不可中途打断。(README)
--pads 调下巴、--resize_factor 降分辨率),身体/姿态完全不动。→ 能"保留"半身/全身构图,但不具备身体动作生成能力。手部/身体穿帮是该类方法的通病(README 的 Tips 里已在教用户用 --nosmooth 处理"two mouths"等伪影)。requirements.txt 逐字(raw):librosa==0.7.0
numpy==1.17.1
opencv-contrib-python>=4.2.0.34
opencv-python==4.1.0.25
torch==1.1.0
torchvision==0.3.0
tqdm==4.45.0
numba==0.48
PyPI 硬证据(curl https://pypi.org/pypi/torch/1.1.0/json):torch==1.1.0 全仓仅 9 个 wheel,上传时间 2019-04-30,Python 标签只有 cp27 / cp35 / cp36 / cp37——没有任何 cp38+ 轮子,文件名为 torch-1.1.0-cp37-cp37m-manylinux1_x86_64.whl(manylinux1 是 2010 年 ABI 基线)。同批老矩阵:torchvision==0.3.0(2019-05-22)、numpy==1.17.1(2019-08-27)、numba==0.48(2020-06-30)。 三层含义:① 装不上 —— Ubuntu 24.04 是 Python 3.12,torch 1.1.0 无对应 wheel,pip install -r requirements.txt 直接失败;② 跑不了 —— 1.1.0 的 CUDA 轮子属 CUDA 9/10 时代,arch 上限远低于 sm_120,必报 no kernel image available;③ 生态全断 —— numba 0.48 / numpy 1.17.1 / librosa 0.7.0 在 Python 3.12 上均无可用轮子。 → 判定:该 requirements.txt 在 2026 年 Ubuntu + sm_120 上属于"不可复现文档",只能当算法参考,必须用现代 torch 重写推理脚本(社区各 fork / ComfyUI 节点本质都是这么做的)。这也解释了为何 Issue 追踪器自 2024 初就没人回答环境问题。
Python 3.6)→ 与上面的 2019 pin 互相印证,是2019 年的运行环境,在 Ubuntu 24.04 + CUDA 13.2 上无法照原样复现。(详细 pin→sm_120 推导见第 5 节。)face_detection/detection/sfd/s3fd.pth,来源是 adrianbulat.com(README 自己提供了备用 SharePoint 链接,说明主链常挂);模型权重走 Google Drive。| 项目 | 数值 | 证据 |
|---|---|---|
| Stars / Forks | 6090 / 981 | ungh.cc/repos/bytedance/LatentSync;GitHub Issues 页 aria-label 亦为 "6090 users starred this repository" |
| 创建 / 最后 push | 2024-12-11 / 2025-06-20 | 同上 → 2026 年零代码更新(截至 2026-09-22 已停更约 15 个月) |
| GitHub Releases | 空({"releases":[]}) | 同上。权重走 HuggingFace,不打 GitHub Release |
| 最新版本 | 1.6(2025/06/11) | README Updates:"2025/06/11: We released LatentSync 1.6, which is trained on 512×512 resolution videos to mitigate the blurriness problem.";HF 仓 ByteDance/LatentSync-1.6 |
"2026 后继版是否改变要求" 的答案:没有 2026 新版;而 1.6 相对 1.5 把显存要求改坏了(8GB → 18GB)。
【作者自述,README 原文,最权威】
Minimum VRAM for inference:
- 8 GB with LatentSync 1.5
- 18 GB with LatentSync 1.6
→ 16GB 单卡 < 官方 18GB 门槛,官方口径下 1.6 不可行。 要留在 16GB 内,官方口径只允许退回 1.5(8GB)。
训练档(同一 README,供参考):stage1.yaml 23GB;stage2.yaml 30GB;stage2_efficient.yaml 20GB(作者注:"suitable for users with consumer-grade GPUs, such as the RTX 3090");stage1_512.yaml 30GB;stage2_512.yaml 55GB。(README)
fp16/权重体积(第三方):ComfyUI 封装仓 README 列出 latentsync_unet.pt (~5GB)、stable_syncnet.pt (~1.6GB),并称 "Reduced VRAM Requirements: Optimized to run on 20GB VRAM (RTX 3090 compatible)"。(ShmuelRonen/ComfyUI-LatentSyncWrapper)⚠️ 该仓为第三方封装,且"20GB / RTX 3090"表述自相矛盾(3090 是 24GB),仅作量级参考。
🟢 决定性第三方实测:RTX 5070 Ti 16GB 真的跑起来了(2026 年 7 月) Issue #365 "TensorRT-RTX optimized inference path and reduced-step speed benchmarks",opened on Jul 3, 2026,作者 Petrus Vermaak(自述 "This work was directed, tested, and validated by Petrus Vermaak... OpenAI Codex served as the implementation and engineering agent")。测试系统:RTX 5070 Ti 16 GB test system —— 同为我们关心的显存档位与 Blackwell 世代。基准夹具原文:
Fixture: official demo video/audio, 9.68s output, 242 frames, 512 face processing, 20 steps, guidance 1.5.
| Backend | Wall seconds | Sec/output second | Change vs exact20 | Peak total VRAM |
|---|---|---|---|---|
| exact20 PyTorch eager | 397.512 | 41.065 | baseline | 15807 MB |
| optimized20 TensorRT-RTX exact | 268.160 | 27.702 | 32.54% faster | 12551 MB |
(逐字引用表格数字;链接:issue #365)
→ 【第三方实测】16GB 单卡在 512×512 / 20 步下可以跑完 LatentSync 1.6:裸 PyTorch 峰值 15807 MB(已占满 16GB 的 96%,极限)、TensorRT-RTX 优化后 12551 MB(留有余量)。官方说的 18GB 是保守门槛,实测 16GB 能压进去,但 PyTorch 裸跑基本贴着天花板。
用户侧旁证(都指向 16GB 不够):
inference_steps [20-50]、guidance_scale [1.0-3.0] 两个可调项(步数越高越慢)。optimized20 779.70s、optimized12 609.50s、optimized8 502.50s、optimized4 384.70s;作者自己警告 "optimized4 was visually plausible in this test but showed the largest objective drift, so I would not present it as equivalent quality"(降步数是"近似"模式,质量会漂)。Engine cosine vs PyTorch 0.9999949036、Full-frame SSIM mean 0.981501、SyncNet confidence 8.344(baseline 8.578)。不支持,且被 Issue 明确暴露。 架构是整段扩散去噪(20 步),没有 chunk 级流水线;Issue #326 标题即为 "How to run inference on longer video?"(issue #326)→ 长视频要自己切段处理。无流式输出、无打断机制。
🔴 最硬的坑:官方依赖锁死在 cu121,不含 sm_120 内核。 requirements.txt 逐字(raw):
torch==2.5.1
torchvision==0.20.1
--extra-index-url https://download.pytorch.org/whl/cu121
diffusers==0.32.2
transformers==4.48.0
decord==0.6.0
accelerate==0.26.1
einops==0.7.0
omegaconf==2.3.0
opencv-python==4.9.0.80
mediapipe==0.10.11
python_speech_features==0.6
librosa==0.10.1
scenedetect==0.6.1
ffmpeg-python==0.2.0
imageio==2.31.1
imageio-ffmpeg==0.5.1
lpips==0.1.4
face-alignment==1.4.1
gradio==5.24.0
huggingface-hub==0.30.2
numpy==1.26.4
kornia==0.8.0
insightface==0.7.3
onnxruntime-gpu==1.21.0
DeepCache==0.1.1
→ torch==2.5.1 + cu121 官方 wheel 不含 sm_120(Blackwell)内核。在 RTX 5060 Ti 上照原样装,会撞上与 Duix #624 同型的 no kernel image is available for execution on the device。必须自行升级到 torch ≥ 2.7 + cu128/cu13x,而任何 torch 升级都可能打破下面这批老旧 pin。这是本项目在 sm_120 上最大的工作量来源。
✅ 好消息:不需要 xformers,也不需要 flash-attn。 我逐条核对了上面的 requirements:没有 xformers,也没有 flash-attn。→ 不存在 flash-attn/triton 需要现场编译 sm_120 内核的问题,这一点比很多扩散类 talking-head 项目(如 MuseTalk / 部分 ComfyUI 节点)友好得多。注意力走 diffusers 默认路径。 (若走 TensorRT-RTX 路线,Issue #365 已证明在 16GB Blackwell 上可行且比 PyTorch 省 3256MB 峰值显存、快 32.54%,是本项目在 16GB 卡上最值得走的优化路径。)
其他版本坑(逐条):
onnxruntime-gpu==1.21.0:需要与 CUDA 版本匹配的 ORT 构建;这类 pin 在 CUDA 13.x 宿主上常需换轮子。insightface==0.7.3 + face-alignment==1.4.1:人脸检测依赖,要现场编译(insightface 需 C++ 工具链),是环境搭建的常见失败点。numpy==1.26.4 + gradio==5.24.0 + mediapipe==0.10.11:整体是 2025 年上半年的老矩阵,与 torch≥2.7 的组合未被作者验证过(作者 2025-06-20 后就停更了)。DeepCache==0.1.1:推理加速用,可能与新版 diffusers 冲突(diffusers 被 pin 在 0.32.2)。setup_env.sh:README 要求 source setup_env.sh 来装依赖+下权重,但我抓取该文件返回空(未核实其内容) —— 不排除 404 或抓取失败,此项标注为未核实。能跑,但要打破官方依赖并接受"离线批处理"定位:官方 18GB 门槛把 1.6 判为不合规;实测 16GB(5070 Ti)在 512×512/20 步下峰值 15807MB 可通过,TensorRT-RTX 降到 12551MB。代价是速度仅约实时的 1/28~1/41,且必须自己把 torch 从 cu121 迁到 cu128+ 才能吃 sm_120。 它不是"实时数字人",是"离线 lip-sync 渲染器"。
| 维度 | Duix.Avatar(原 Duix.Heygem) | Wav2Lip | LatentSync |
|---|---|---|---|
| 2026 状态 | 改名后仍活跃(2026-04 有提交) | 停更(2025-06-22 最后一次 push,README 已商业导流) | 停更 15 个月(2025-06-20) |
| Stars | 15547 | 13219 | 6090 |
| 官方推理显存 | 未给数字(整机推荐 RTX 4070,16GB+ 为推荐档) | 未给 | 8GB(1.5) / 18GB(1.6) |
| 16GB 单卡 | ⚠️ 怕是不行:三容器常驻,无 16GB 实测 | ✅ 几无悬念可跑(推断) | ⚠️ 官方不合规,实测 15807MB 压线可跑 |
| 成片速度 | 未查到 FPS(官方定义非实时) | 未查到(Issue 普遍抱怨慢) | 27.7~41.1 秒 / 1 秒成片(5070Ti 16GB 实测) |
| 流式/打断 | ❌ 提交+轮询 | ❌ 整段批处理 | ❌ 整段扩散 |
| 中文 | ✅ 8 语言含中文 | ⚠️ 语言无关,但只训练英文(LRS2) | ✅ 明确优化中文视频 |
| 半身/全身 | 【推断】继承源视频(face2face) | 只改嘴部,身体不动 | ❌ 仅面部 |
| sm_120 现状 | 🔴 坏(Issue #624 Open 零回复,主线镜像无 CC12.0 内核;仅 5090 专用镜像存在,未验证 5060 Ti) | 🔴 pin 即不可用(torch==1.1.0 为 2019 年 CUDA 9/10 时代,且无 cp38+ 轮子;须整套重写) | 🟡 需自行把 cu121→cu128+(无 xformers/flash-attn,省一大坑) |
| 商用许可 | ⚠️ 自定义许可(非开源):DUIX.COM COMMUNITY LICENSE,MAU > 1000 即须申请商业授权(README 误写"10 万",差 100 倍) | ❌ 严格禁止商用(LRS2 数据集限制) | ✅ Apache License 2.0(LICENSE 正文已核实) |
| 部署形态 | Docker 三容器 + 客户端 App | 裸 Python 脚本 | 裸 Python 脚本 |
【官方,PyTorch 官方博客】PyTorch 2.7 发布说明逐字(pytorch.org/blog/pytorch-2.7,180248 字节,我已抓取正文):
pip install torch==2.7.0 --index-url https://download.pytorch.org/whl/cu128→ 基线结论:要原生吃 sm_120,torch 必须 ≥ 2.7 且用 cu128(或更新)轮子;cu121 及更早的官方轮子不含 sm_120 内核。
【第三方实测,逐字警告文本 —— 索引标题级证据】 Comfy-Org/ComfyUI Issue #7127 的页面标题逐字包含运行期警告:
UserWarning: NVIDIA GeForce RTX 5070 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation. The current PyTorch install supports CUDA capabilities sm_50 sm_60 sm_61 sm_70 sm_75 sm_80 sm_86 sm_90.
→ 这就是 cu121 轮子的实际 arch 列表:最高到 sm_90,没有 sm_120。触发设备是 RTX 5070 Ti(同为 sm_120),与我们的 5060 Ti 同代同 cap。 ⚠️ 诚实标注:本轮 github.com 的 Issue 正文页多次抓取超时(网络问题),上述引文来自搜索索引返回的页面标题原文,我未读取该 Issue 正文;pytorch/pytorch #166794 同理,其标题逐字为 "[Bug] RTX 5070 Ti (sm_120) not recognized by PyTorch 2.5.1+cu121"。两者作为"同型 pin 已被报告失败"的证据成立,但均为标题级、非正文级证据。 【第三方实测,同型】HybridRobotics/Berkeley-Humanoid-Lite Issue #49 标题逐字:"RTX 5090 (SM_120) fails with default torch/cu121: no kernel image is available"。
| 仓库 | pin(逐字) | sm_120 可用性 | 必须做什么 |
|---|---|---|---|
| LatentSync | torch==2.5.1 / torchvision==0.20.1 / --extra-index-url .../cu121 / onnxruntime-gpu==1.21.0 | 🔴 不可用:2.5.1 早于 2.7,cu121 arch 表止于 sm_90 | 升 torch ≥2.7 + cu128,并连带处理 5.4 的脆点 |
| Wav2Lip | torch==1.1.0 / torchvision==0.3.0 / numba==0.48 / numpy==1.17.1 / librosa==0.7.0 | 🔴 绝对不可用:2019 年 CUDA 9/10 产物,且无 cp38+ 轮子 | pin 形同废弃,须整体现代化重写 |
| Duix.Avatar | 源码无 pin——pin 在 Docker 镜像内部 | 🔴 已实测失败:Issue #624 证实镜像不含 CC 12.0 内核 | 用户无法自行解决,只能等官方重建镜像或自建镜像 |
特别注意 LatentSync 与 #166794 的 pin 是同型:torch 2.5.1 + cu121 正是被反复报告在 RTX 5070 Ti / 5090(sm_120) 上失败的那一组。这不是理论风险,而是已在同代硬件上被复现的失败组合。
torch==1.1.0 到底有多致命(本轮新查,PyPI 硬证据)【第三方可核实,PyPI JSON API】torch==1.1.0 的全部 9 个 wheel(https://pypi.org/pypi/torch/1.1.0/json):
torchvision==0.3.0 2019-05-22、numpy==1.17.1 2019-08-27、numba==0.48 2020-06-30)cp27 / cp35 / cp36 / cp37 —— 没有任何 cp38+ 轮子,也没有 aarch64;torch-1.1.0-cp37-cp37m-manylinux1_x86_64.whl(manylinux1 = 2010 年 ABI 基线)requires_python 字段为空(2019 年的包普遍不声明)三层含义:
pip install -r requirements.txt 必然失败;若强行从源码编译,2019 年的 CUDA 代码在 CUDA 13.2 的 nvcc 下几乎必然编译失败。no kernel image is available for execution on the device(与 Duix #624 同型错误)。numba==0.48(依赖 llvmlite 0.31)、numpy==1.17.1、librosa==0.7.0 在 Python 3.12 上均无可用轮子。→ 最终判定:Wav2Lip 的 requirements.txt 在 2026 年的 Ubuntu + sm_120 上属于"不可复现文档"。它无法按官方方式部署,只能把 wav2lip.pth 权重 + 推理逻辑搬到现代 torch 上重写(社区各种 fork 与 ComfyUI 节点本质都是这么做的)。这与 README 自述的 Python 3.6 前提互相印证,也解释了为何其 Issue 追踪器自 2024 年初起就无人处理环境类问题。
唯一的硬门槛是 torch 2.5.1 → ≥2.7(+cu128),但换版本号会牵动一串 2025 年上半年的旧 pin:
decord==0.6.0:视频解码轮子,长期只发布到 cp310/cp311 附近,新 Python 上常无轮子 → 这类项目的经典断点。insightface==0.7.3:需现场 C++ 编译(人脸检测依赖),是环境搭建的常见失败点。DeepCache==0.1.1 + diffusers==0.32.2:DeepCache 是推理加速补丁,与新版 diffusers 的组合从未被作者验证(作者 2025-06-20 后停更)。onnxruntime-gpu==1.21.0:与你的 CUDA 13.2 宿主最容易打架,单独见 5.5。✅ 一个重要的好消息:LatentSync 的 requirements.txt 没有 xformers,也没有 flash-attn → 不需要为 sm_120 现场编译 flash-attn / triton 内核。中文资料里常见的那类报错(如 cnblogs 记录的 "nvcc fatal : Unsupported gpu architecture 'compute_120'")在 LatentSync 的原生路径上不会出现,只在你想额外加装 flash-attn 时才会撞上。若走 Triton 路线,PyTorch 2.7 已内置 Triton 3.3,官方称其 "adds support for the Blackwell architecture with torch.compile compatibility"。
onnxruntime-gpu==1.21.0 与 CUDA 13.2 宿主的冲突(本轮新查)【官方,ONNX Runtime CUDA EP 文档逐字】(onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider,正文已抓取):
→ onnxruntime-gpu==1.21.0 是 CUDA 12.8 + cuDNN 9.x 构建,而你的宿主是 CUDA 13.2(跨大版本)。按 NVIDIA Minor Version Compatibility,12.8 构建只保证与 12.x 兼容。 → 【推断,风险项】两条路:① 额外提供 CUDA 12 运行库(如 pip 的 nvidia-cuda-runtime-cu12 / nvidia-cudnn-cu12),否则可能缺 libcudart.so.12;② 把 ORT 升到 ≥1.27(cu13.0 构建)以匹配宿主 —— 但后者已脱离 LatentSync 的 pin,必须重测。 → Blackwell 专门适配:搜索索引显示 microsoft/onnxruntime 有一个 PR #23928 "Extend CMAKE_CUDA_FLAGS with all Blackwell compute capacity"(标题逐字),说明 ORT 曾专门为 Blackwell 补齐 CUDA arch 编译标志。⚠️ 该 PR 进入哪个发布版本我未能核实(PR 页本轮网络不可达),标注为未核实;实务上 cu12.8 构建通常可借 PTX JIT 在 sm_120 上跑通,但需实测确认。
| 仓库 | 16GB 上能否跑 | 改动量 | 主要拦路虎 |
|---|---|---|---|
| LatentSync | ✅ 实测可行(峰值 15807MB 压线 / TensorRT 后 12551MB) | 中:换 torch≥2.7+cu128,处理 decord / ORT / DeepCache | 无 flash-attn 坑是最大优势;ORT 与 CUDA 13.2 需调和 |
| Duix.Avatar | ❌ 目前不可行(#624 已实测同型失败) | 无法自行解决(内核编译在官方镜像里) | 官方 #624 零响应;仅 5090 专用镜像,未验证 5060 Ti |
| Wav2Lip | ⚠️ 理论上限低但无法按官方方式部署 | 大:2019 依赖整套作废,须重写推理脚本 | torch==1.1.0 无 cp38+ 轮子;且严格禁止商用 |
一句话:三个项目的 pin 无一为 sm_120 准备。LatentSync 只差一个 torch 升级(且没有 flash-attn 这个天坑),是三者中唯一在 16GB Blackwell 上被实测跑通的;Duix 卡在官方镜像、用户无力自救;Wav2Lip 的依赖本身就是 2019 年化石。
Duix 系
no kernel image is available(2026-09-04 开,Open 零回复)Wav2Lip
LatentSync
依赖 pin / sm_120 / 许可 专项(本轮补查)
torch==1.1.0 / torchvision==0.3.0 / numba==0.48 / numpy==1.17.1 / librosa==0.7.0sm_50 ... sm_90(无 sm_120)(索引标题级证据)未查到 / 未能核实的项(诚实标注)
hub.docker.com 本机多次超时) —— 未查到duixcom org 完整仓库列表 —— 官方 API 限流(core remaining=0),改用 ungh 逐仓探测,可能遗漏 org 下其他未探测的仓commits/<branch>.atom 方案本轮多次超时失败,故全文采用 ungh.cc 的 pushedAt(Duix-Avatar 2026-04-21T07:06:36Z、Wav2Lip 2025-06-22T02:41:21Z、LatentSync 2025-06-20T07:36:58Z),该字段为"最后一次 push"而非严格最后一次 commit,量级与结论不受影响github.com HTML 抓取持续超时,仅有页面标题逐字,已逐条标注为"索引标题级证据"hub.docker.com 镜像 tag 体积与更新日期 —— 未查到(多次超时)