# C2 组核查结果：实时数字人（talking head）开源项目 — 3 仓深挖

- **核查日期**：2026-09-22（GitHub 数据通过 ungh.cc 镜像实时读取；GitHub 官方 API 本机 IP 已触发 60 次/时限流，故未使用官方 API 计数）
- **目标硬件**：RTX 5060 Ti 16GB / sm_120 (Blackwell) / 32GB RAM / Ubuntu；驱动 610.43.02，CUDA 13.2
- **标注约定**：`【作者自述】`＝仓库 README / Release 官方口径；`【第三方实测】`＝Issue / 博客 / Wiki 中他人给出的数字；`【推断】`＝基于文件与文档的合理推导，非直接证据；`未查到`＝未能验证，不编造。

---

## 0. 头条结论（先说最重要的三条）

1. **`duixcom/Duix.Heygem` 这个仓库在 2026 年已经不存在了 —— 它被官方重命名为 `duixcom/Duix-Avatar`**，GitHub 对旧地址做 301 跳转。所以"Duix.Heygem 2026 状态"的正确答案是：**主体已改名，改名后的 Duix-Avatar 是唯一活跃主线**。
2. **Duix.Avatar 在 Blackwell / sm_120 上目前是坏的**：2026-09-04 有人开了 Issue #624，报 `CUDA error: no kernel image is available for execution on the device`，根因指向 Docker 镜像里的 PyTorch 不含 CC 12.0 内核，**该 Issue 至今 Open 且零回复**。
3. **LatentSync 官方 1.6 的最低推理显存是 18GB，16GB 单卡不达官方门槛**；但 2026-07-03 有人在 **RTX 5070 Ti 16GB** 上实测跑通了（峰值 15807MB），代价是约 27.7–41.1 秒墙钟时间换 1 秒输出 —— 约为实时的 1/28 ~ 1/41。

---

## 1. duixcom/Duix.Heygem（→ 已改名 duixcom/Duix-Avatar）

### 1.1 仓库身份与 2026 活跃度

| 项目 | 结果 | 证据 |
|---|---|---|
| `duixcom/Duix.Heygem` | **301 跳转到 `duixcom/Duix-Avatar`** | `curl -sL -w '%{url_effective}'` → `200 -> https://github.com/duixcom/Duix-Avatar`；ungh.cc 对 `duixcom/Duix.Heygem` 查询返回 `"id":907627874,"name":"Duix-Avatar"`（同一个 repo id，证明是重命名而非新仓） |
| Stars / Forks | **15547 / 2642** | [ungh.cc/repos/duixcom/Duix-Avatar](https://ungh.cc/repos/duixcom/Duix-Avatar)，2026-09-22 读取 |
| 创建 / 最后 push | 创建 `2024-12-24T03:00:26Z`；**最后 push `2026-04-21T07:06:36Z`** | 同上。→ **2026 年有提交（4 月），但到 9 月已约 5 个月无代码更新** |
| 改名证据（Release） | v1.0.5（2025-08-15）发布说明：**"The original project name "HeyGem" has now been officially changed to "Duix.Avatar"."**；v1.0.6（2025-09-28）：**"It's just a name change. If the version you're using has no issues, you can continue to use it."** | [releases](https://github.com/duixcom/Duix-Avatar/releases) |

### 1.2 `duixcom` org 下各仓 2026 状态（逐个核查，含"是否存在"）

> 说明：任务要求的 `curl https://api.github.com/orgs/duixcom/repos` 因本机 IP **限流（core remaining=0，reset 13:52）**无法执行；改用 ungh.cc 镜像逐个精确查询 + 落地页跳转探测，结论同样可证伪。

| 仓库 | 状态（2026-09-22） | Stars | 最后 push | 说明 |
|---|---|---|---|---|
| `duixcom/Duix-Avatar`（原 Duix.Heygem） | ✅ 存在、2026 有更新 | 15547 | 2026-04-21 | **本组主目标的正身** |
| `duixcom/Duix.Mobile`（原 Duix.mobile） | ✅ 存在、**2026 活跃** | 8252 | **2026-08-05** | 移动端/嵌入式**实时交互** SDK，非本地 GPU 服务。仓描述：**"on-premise deployment and <1.5 s latency"** |
| `duixcom/Duix.Heygem.Android` | ❌ **404，不存在** | — | — | 该仓从未存在或已删除；`raw` 的 main/master 均 404 |
| `duixcom/duix-skills` | ⚠️ **2026-07-14 新建**，但非本地模型 | 4 | 2026-08-12 | 见下方"2026 后继者判定" |
| `duixcom/Duix-Reface` | ❌ **404**（`Duix.Mobile` README 里仍在推荐它，但已不可访问） | — | — | 文档链接已失效 |
| `duixcom/Duix.mobile`（旧名） | 301 → `duixcom/Duix-Mobile` | — | — | 同为重命名 |

**"2026 后继者是否取代 Duix.Heygem" 判定：**
- **取代 Duix.Heygem 的就是 Duix-Avatar 本身**（同一 repo id 的重命名），**不是**另起的新项目。
- `duixcom/duix-skills`（2026-07-14 新建）**不构成本地部署的后继者**：它是一套给 AI agent 用的**云端 API skill**，需要 `DUIX_APP_ID`/`DUIX_APP_KEY`/`DUIX_API_KEY` 凭据，走**订阅套餐 + 预付费积分**计费，产出是 `conversation_url` 网页链接或云端生成的 MP4；其 README 明确写 "Prepaid credits (consumed by task video duration)"、"[Production time] About 20 minutes to 2 hours"。→ **它把 Duix 的 lip-sync 能力 SaaS 化了，与"本地 16GB 显卡部署"是两条路**。（证据：[duix-skills README](https://github.com/duixcom/duix-skills)）

### 1.3 (c) 显存需求 + 16GB 单卡能否跑

- **【作者自述】README 只给整机推荐，未给显存数字**：`CPU: 13th Gen Intel Core i5-13400F / Memory: 32GB / Graphics Card: RTX 4070`，硬盘 C 盘 >100GB、D 盘 >30GB；服务器部署要下载 **约 70GB** 流量。（[README](https://github.com/duixcom/Duix-Avatar)）
- **【第三方实测/整理，AI 生成 Wiki，可信度中】** DeepWiki 的硬件页给出显存表：`GPU | NVIDIA RTX 4070 (8GB VRAM) | NVIDIA RTX 4090 / RTX 5090 (16GB+ VRAM)`（Minimum → Recommended），并给出调优表：`8GB → max_split_size_mb 256/512`、`12-16GB → 512 (default)`、`24GB+ → 1024`。
  ⚠️ **该页有硬错误需警惕**：它把 RTX 5090 的 "CUDA Compute" 标为 **9.0**（5090 实为 sm_120），说明此页为 AI 自动生成、未经人工校对；其数字仅可作参考。链接：[deepwiki.com/duixcom/Duix-Avatar/6.3-hardware-and-performance-tuning](https://deepwiki.com/duixcom/Duix-Avatar/6.3-hardware-and-performance-tuning)
- **【第三方实测，原始 Issue 标题】** Issue #325 标题即为 **"6G VRAM can't run with lite version"** → 6GB 不可行（lite 版）。链接：[issue #325](https://github.com/duixcom/Duix-Avatar/issues/325)
- **融合判断【推断】**：官方推荐 16GB+ 显存即可（4090 24G / 5090 32G 为推荐档），16GB 单卡处在**推荐档下沿**；但因为服务是**三个容器同时常驻**（TTS + ASR + 视频合成，其中 TTS 还要做声音克隆训练），16GB 是否够用**取决于三服务是否同卡共存**——这点**未查到 16GB 单卡的实测报告**，建议列为"高风险待实测"。
- **fp16 权重体积：未查到**。Duix 走 Docker 打包分发，权重内嵌在镜像里（`guiji2025/duix.avatar`），仓库不发布独立 `.pt`/`.safetensors`，故无 fp16 权重体积数字可引用。Docker Hub 我尝试拉取 tag 体积，**多次超时未成功**（`hub.docker.com` 在本机网络下不可达）。

### 1.4 (d) 端到端时延 / FPS

- **【作者自述】Duix.Avatar 本身就不是实时的**，README 原文：**"Duix.Avatar's digital human realizes digital human cloning and non-real-time video synthesis."** 以及 **"If you want a digital human to support interaction, you can visit duix.com"** → 实时交互能力被**刻意排除**在开源版之外。（[README](https://github.com/duixcom/Duix-Avatar)）
- **具体 FPS / 首帧时延：【未查到】**（README 与 Release 均无数字）。
- **反证（说明性能问题真实存在）**：Issue #251 标题 **"调优，咨询一下，我电脑上有4张24G的显卡，推理速度很慢，如何提升速度"** → 用户持 4×24GB 仍嫌慢。（[issue #251](https://github.com/duixcom/Duix-Avatar/issues/251)）
- **注意区分**：`Duix.Mobile` 的时延数字**不能**用于 Duix.Avatar —— Mobile README 写 **"AI avatar response latency under 120ms (tested on Snapdragon® 8 Gen 2 SoC)"**，那是手机 SoC 上的端侧渲染，与 GPU 服务的视频合成是两套东西。（[Duix-Mobile README](https://github.com/duixcom/Duix-Mobile)）

### 1.5 (e) 流式 / 可打断

- **Duix.Avatar：不支持**。API 形态是**提交+轮询**，非流式：合成接口 `http://127.0.0.1:8383/easy/submit`，进度查询 `http://127.0.0.1:8383/easy/query?code=${taskCode}`（GET 轮询）；TTS 接口参数里甚至有 `"streaming": false, // Fixed parameter` 这一固定值。→ **非流式、不可打断**。（[README](https://github.com/duixcom/Duix-Avatar)）
- **Duix.Mobile：支持流式 + 打断（barge-in）**。README FAQ 原文：**"Yes, streaming audio with barge-in support is available from the July 17, 2025 release."**，另一处列出 **"Streaming Audio Support: Synthesize and speak simultaneously, supports interruption and barge-in"**。（[Duix-Mobile README](https://github.com/duixcom/Duix-Mobile)）

### 1.6 (f) 中文支持 / 半身全身

- **中文：支持**。【作者自述】README 原文：**"Multi-language Support: Scripts support eight languages - English, Japanese, Korean, Chinese, French, German, Arabic, and Spanish."**（含中文，共 8 种）（[README](https://github.com/duixcom/Duix-Avatar)）
- **半身/全身【推断】**：技术路线是 **face2face（换脸式口型驱动）**——视频合成容器把数据目录挂成 `~/duix_avatar_data/face2face:/code/data`，即**在源视频上重绘人脸区域**。据此推断**输出构图继承源视频**，源视频是半身/全身就得到半身/全身；但这属于**推断，非文档明示**，我未查到官方对"半身/全身"的明确能力声明。证据：[deploy/docker-compose-linux.yml](https://raw.githubusercontent.com/duixcom/Duix-Avatar/main/deploy/docker-compose-linux.yml)

### 1.7 (g) 坑：sm_120 / Blackwell / Docker 部署 / 许可

#### 🔴 最致命：Blackwell 目前跑不起来（2026 年 9 月仍是 Open）
Issue **#624**，标题 **"RTX 5070 (Blackwell) not supported — CUDA kernel error on Docker backend"**，**opened on Sep 4, 2026**，状态 **Open，且页面零回复**（"Sign up for free to join this conversation"）。原文逐字引用：

> GPU: NVIDIA GeForce RTX 5070 Laptop GPU
> Compute Capability: 12.0 (Blackwell)
> Driver: 592.15
> Error from duix-avatar-tts container:
> `RuntimeError: CUDA error: no kernel image is available for execution on the device`
> Root cause: The Docker images are compiled with a PyTorch version that doesn't include CUDA kernels for compute capability 12.0. PyTorch 2.7+ added Blackwell support.
> Request: Please rebuild and push updated Docker images (`guiji2025/fish-speech-ziming`, `guiji2025/fun-asr`, `guiji2025/duix.avatar`) compiled against PyTorch 2.7+ with sm_120 support.

链接：[issue #624](https://github.com/duixcom/Duix-Avatar/issues/624)

→ **对 RTX 5060 Ti (sm_120) 的直接含义：默认三个 Docker 镜像的主线编译目标不含 sm_120，TTS 容器会直接抛 `no kernel image is available`。维护者截至 2026-09-22 未回应。**

#### ⚠️ 50 系有独立镜像，但只针对 5090，且未验证 sm_120/5060 Ti
- README What's New 原文：**"[Nvidia 50 Series GPU Version Notice] 1. Tested and verified on 5090 GPU"**；服务器部署章节原文：**"For 50 series graphics cards (tested and also works for 30/40 series with CUDA 12.8) Uses the official preview version of PyTorch"**。（[README](https://github.com/duixcom/Duix-Avatar)）
- 确实存在单独的 50 系 compose：`deploy/docker-compose-5090.yml`（HTTP 200，1064 字节），镜像换成了 **`guiji2025/duix.avatar-5090`** 与 **`guiji2025/fish-speech-5090`**，仍保留 `PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512` 与 `shm_size: '8g'`。
  → 【推断】**50 系走单独镜像这条路径"理论上"是给 Blackwell 的**，但官方只声明在 **5090** 上验证过；**在 5060 Ti 上是否可用未查到任何证据**，且 #624 的存在说明普通镜像确实没有 sm_120 内核。**这是 16GB 选型里最大的未知风险点。**
  （证据：[docker-compose-5090.yml](https://raw.githubusercontent.com/duixcom/Duix-Avatar/main/deploy/docker-compose-5090.yml)）

#### 部署形态：纯 Docker 打包，无源码级可调
三个服务（[docker-compose-linux.yml](https://raw.githubusercontent.com/duixcom/Duix-Avatar/main/deploy/docker-compose-linux.yml)）：

| 服务 | 镜像 | 端口 | 关键配置 |
|---|---|---|---|
| duix-avatar-tts | `guiji2025/fish-speech-ziming` | 18180:8080 | `runtime: nvidia`，`NVIDIA_VISIBLE_DEVICES=0` |
| duix-avatar-asr | `guiji2025/fun-asr` | 10095:10095 | `runtime: nvidia`，**`privileged: true`** |
| duix-avatar-gen-video | `guiji2025/duix.avatar` | 8383:8383 | `runtime: nvidia`，**`privileged: true`**，`shm_size: '8g'`，`PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512` |

**坑点**：
- 全部容器 `runtime: nvidia`，其中两个 **`privileged: true`** → 与宿主驱动强耦合；**README 明确 NVIDIA 显卡与驱动是硬性前提**，且 FAQ 原文 **"All computing power of this project is local. The three services won't start without an NVIDIA graphics card or proper drivers."**
- **必须装 NVIDIA Container Toolkit** 并 `sudo nvidia-ctk runtime configure --runtime=docker`（README 给了完整步骤）；**没有 CPU 回退路径**（DeepWiki：**"The system has no CPU-only fallback; an NVIDIA GPU is mandatory for all deployment modes."**）
- `shm_size: '8g'` 是硬需求：DeepWiki 把 **"Bus Error (Core Dumped)"** 归因于 **"Insufficient shared memory → Increase shm_size to 12g or 16g"**；1080p/60s 就建议 8GB，4K 建议 12GB+。
- **磁盘/流量重**：安装下载 **约 70GB 流量**，**约半小时**（README）；C 盘 >100GB。32GB RAM 是 README 的最低推荐（Ubuntu 版原文写 **"Memory: 32G or more (necessary)"**）—— 我们的 32GB 正好卡线。
- **Ubuntu 支持面窄**：**"We have conducted a complete test on Ubuntu 22.04"**，内核 **6.8.0-52-generic** 验证过，**其他 Linux 版本未做兼容性测试**。若目标是 Ubuntu 24.04，属于未验证组合。

#### 许可 / 商业授权 —— ⚠️ README 与 LICENSE 正文数字不一致，以 LICENSE 为准

**我在本轮补抓了 LICENSE 正文**（`raw.githubusercontent.com/duixcom/Duix-Avatar/main/LICENSE`，7120 字节），**标题逐字为 `DUIX.COM COMMUNITY LICENSE AGREEMENT`**（自定义社区许可，非标准 SPDX 许可）。

**🔴 关键纠正：商用门槛是 1000 MAU，不是 README 说的 10 万用户。**
- 【作者自述，README 英文对比表】原文：**"Commercial Authorization: Supports global free commercial use (enterprises with more than 100,000 users or annual revenue exceeding 10 million USD need to sign a commercial license agreement)"** → 白话是"免费商用，>10 万用户或年营收 >1000 万美元才需签约"。
- 【法律文本，LICENSE 第 2 条逐字】**"2. Additional Commercial Terms. If, on the DUIX.COM version release date, either (a) the Monthly Active Users of the products or services made available by or for Licensee, or Licensee's affiliates, is greater than 1 thousand in the preceding calendar month, or (b) your product incorporating DUIX.COM Materials has greater than 1 thousand Monthly Active Users, you must request a commercial license from DUIX.COM, which DUIX.COM may grant to you in its sole discretion, and you are not authorized to exercise any of the rights under this Agreement unless or until DUIX.COM otherwise expressly grants you such rights."**
- 且 LICENSE 给出了 **"Monthly Active Users" 的定义**逐字：**"the number of unique users who interact with your product or service that incorporates the DUIX.COM Materials at least once during a calendar month."**
→ **两者相差 100 倍**（1000 vs 100,000）。**以 LICENSE 正文为准：月活超过 1000 就必须申请商业授权，且是否授予由 DUIX 单方裁量；未获授权前不得行使本协议任何权利。** 这对"做个小产品上线"几乎是立刻触发的门槛，**强烈建议报告里按 1000 MAU 写，并注明 README 存在 100 倍口径差异**。

**LICENSE 里的其他附加义务（逐字摘要，易被忽略但都是硬性）**：
- **显著署名**：第 1.b.i 条要求 (B) **prominently display "Built with DUIX.COM"** on a related website / user interface / blogpost / about page / product documentation（须在网站、UI、博客、关于页或产品文档显著位置展示"Built with DUIX.COM"）。
- **服务条款内声明**：(C) 须在 ToS / EULA 中**清晰声明**你的产品包含或基于 DUIX.COM 技术。
- **AI 模型命名强制前缀**：若用其材料训练/微调模型并分发，**"you shall also include \"DUIX.COM\" at the beginning of any such AI model name"**。
- **Notice 文件**：分发时须保留 `Notice` 文本文件，含 **"DUIX.COM is licensed under the DUIX.COM Community License, Copyright © DUIX.COM Platforms, Inc. All Rights Reserved."**
- **案例授权（第 5.c 条）**：使用即视为授予 DUIX **永久、全球、非独占、免版税、不可撤销**的许可，可用你的案例/示例做营销与产品改进，**"without any compensation or attribution to you"**。
- **专利/诉讼终止条款（第 5.b 条）**：你若对 DUIX 提起侵权诉讼，许可**自起诉之日起自动终止**。

（对照：**bytedance/LatentSync 的 LICENSE 正文我已抓取核实，为标准 `Apache License Version 2.0`**，详见 3.6 节。）

同时 README 明说自部署版的定位是 **"Lip Sync Effect: Usable effect"（可用级）**，而云端 API 才是 **"Stunning and higher definition effect"**、且 **"Iteration Speed: Slow updates, bug fixes depend on the community"** —— 官方自认开源版迭代慢。

---

## 2. Rudrabha/Wav2Lip

### 2.1 2026 状态：**事实性停更（effectively abandoned）**

| 项目 | 数值 | 证据 |
|---|---|---|
| Stars / Forks | **13219 / 2849** | [ungh.cc/repos/Rudrabha/Wav2Lip](https://ungh.cc/repos/Rudrabha/Wav2Lip) |
| 创建 | 2020-08-07 | 同上 |
| **最后 push** | **2025-06-22** | 同上。→ 2026 年**零更新** |
| 最新 Issue 活动 | 搜索页里最新的 Issue 是 **#613 "Extracting raw audio... and stuck. Load not working"，opened on Jan 3, 2024**；#584 **opened on Nov 10, 2023** | [issues 搜索](https://github.com/Rudrabha/Wav2Lip/issues?q=is%3Aissue+VRAM) |

**停更的硬证据（README 已被改成商业导流页）**：`master` 分支 README **开头第一段**不再是论文说明，而是：

> # Commercial Version
> Create your first lipsync generation in minutes. Please note, **the commercial version is of a much higher quality than the old open source model!**
> Create your API key from the [Dashboard](https://sync.so/keys).

其后大段内容是 Sync.so 的 `pip install syncsdk` 调用示例，论文原文被挤到 README **尾部**。（[README](https://github.com/Rudrabha/Wav2Lip)）
→ 结合 **2025-06-22 那次 push 的性质**（把 README 换成商业 API 导流，而非代码更新），以及 **Issue 追踪器自 2024 年初起基本无维护者回应**，可以判定：**Wav2Lip 在 2026 年已是"存档态"项目，作者精力已转移到商业产品 sync.so。**

**社区推荐的后继者**：
- 【可核实】**LatentSync（bytedance）** 的 README 明确把自己的血统写进了致谢：**"Some code are borrowed from MuseTalk, StyleSync, SyncNet, Wav2Lip."** → 后继路线在开源侧就是 **LatentSync / MuseTalk** 这一支扩散模型方案。（[LatentSync README](https://github.com/bytedance/LatentSync)）
- 【可核实】**作者本人的商业后继是 sync.so**（README 与 LICENSE 章节均指向 `rudrabha@synclabs.so` / `prajwal@synclabs.so` 与 `https://synclabs.so/`）。
- ⚠️ "社区公认的替代品排序"属于主观结论，我**未查到**一份权威的二选一推荐，故此处只给可核实的血缘关系，不替社区下结论。

### 2.2 (c) 显存需求 + 16GB

- **【作者自述】README 完全没有给显存要求** —— 连"Recommended GPU"都没有。只给了软件前提 `Python 3.6`、`ffmpeg`。（[README](https://github.com/Rudrabha/Wav2Lip)）
- **具体 fp16 权重体积 / 实测峰值显存：未查到。** 官方权重走 **Google Drive** 分发（`Wav2Lip` 与 `Wav2Lip + GAN` 两个 checkpoint），仓库内不发布文件；我尝试检索 CSDN/Reddit 的实测数字，**未获得可引用的具体数值**，故不编造。可确认的只有量级判断：
- **【推断】**Wav2Lip 主干是 96×96 输入的小型卷积网络（判别器 SyncNet 同量级），**模型本身显存需求远低于扩散类方案**；真正的开销来自人脸检测（S3FD）与逐帧 IO。**在 16GB 卡上"能跑"几乎没有悬念**，瓶颈是速度与画质而非显存。⚠️ 此段为**推断**，非实测。

### 2.3 (d) 端到端时延 / FPS

**未查到官方或可引用的实测数字。** 唯一相关线索是 Issue **#584**（标题 **"How much time do you need to lip sync a 10 sec or 1 minute video?"**，opened on **Nov 10, 2023**，**至今 Open 且无人给出有效回答**），提问者原文：

> I have been trying the last days with both wav2lip HD (not in auto) and retalker, and found that **both are slow and very GPU consuming.**
> I would like to know everyone of you HOW MUCH GPU do you use (what card) and HOW MUCH time does it take for you to do it?
> Please contribute. Because **I am about to drop this technology and give up on it**, maybe others peoples experiences will give me hope.

链接：[issue #584](https://github.com/Rudrabha/Wav2Lip/issues/584)
→ 该 Issue **本身没有给出数字**，它只能证明"用户普遍抱怨慢、且社区无人应答"。

### 2.4 (e) 流式 / 可打断

**不支持。** 架构是整段批处理：`python inference.py --checkpoint_path <ckpt> --face <video.mp4> --audio <an-audio-source>`，结果一次性写盘为 `results/result_voice.mp4`。**无流式接口、无 chunk 级回调、不可中途打断。**（[README](https://github.com/Rudrabha/Wav2Lip)）

### 2.5 (f) 中文支持 / 半身全身

- **中文：语言无关，可用但非优化项。**【作者自述】README 原文：**"Works for any identity, voice, and language. Also works for CGI faces and synthetic voices."** → 因为它是**纯音频驱动**（mel 频谱 → 口型），语言不进入模型，**中文可跑**。⚠️ 但训练集是 **LRS2（英文）**，README 的免责声明原文明说 **"As the models are trained on the LRS2 dataset"** → 中文口型准确度**无官方保证**。【推断】中文效果弱于英文。
- **半身/全身：不驱动身体，只改嘴部区域。** 模型在检测到的人脸框内工作（建议用 `--pads` 调下巴、`--resize_factor` 降分辨率），**身体/姿态完全不动**。→ **能"保留"半身/全身构图，但不具备身体动作生成能力**。手部/身体穿帮是该类方法的通病（README 的 Tips 里已在教用户用 `--nosmooth` 处理"two mouths"等伪影）。

### 2.6 (g) 坑：sm_120 / 依赖 / 维持性

- **🔴 依赖 pin 是 2019 年的，对 sm_120 属于"绝对不可用"**（本轮补抓，纠正我上一版的"无版本锁定"说法）。`requirements.txt` 逐字（[raw](https://raw.githubusercontent.com/Rudrabha/Wav2Lip/master/requirements.txt)）：

```
librosa==0.7.0
numpy==1.17.1
opencv-contrib-python>=4.2.0.34
opencv-python==4.1.0.25
torch==1.1.0
torchvision==0.3.0
tqdm==4.45.0
numba==0.48
```

  **PyPI 硬证据**（`curl https://pypi.org/pypi/torch/1.1.0/json`）：`torch==1.1.0` 全仓**仅 9 个 wheel**，上传时间 **2019-04-30**，Python 标签**只有 cp27 / cp35 / cp36 / cp37——没有任何 cp38+ 轮子**，文件名为 `torch-1.1.0-cp37-cp37m-manylinux1_x86_64.whl`（manylinux1 是 2010 年 ABI 基线）。同批老矩阵：`torchvision==0.3.0`（2019-05-22）、`numpy==1.17.1`（2019-08-27）、`numba==0.48`（2020-06-30）。
  **三层含义**：① **装不上** —— Ubuntu 24.04 是 Python 3.12，torch 1.1.0 无对应 wheel，`pip install -r requirements.txt` 直接失败；② **跑不了** —— 1.1.0 的 CUDA 轮子属 CUDA 9/10 时代，arch 上限远低于 sm_120，必报 `no kernel image available`；③ **生态全断** —— numba 0.48 / numpy 1.17.1 / librosa 0.7.0 在 Python 3.12 上均无可用轮子。
  → **判定：该 requirements.txt 在 2026 年 Ubuntu + sm_120 上属于"不可复现文档"，只能当算法参考，必须用现代 torch 重写推理脚本**（社区各 fork / ComfyUI 节点本质都是这么做的）。这也解释了为何 Issue 追踪器自 2024 初就没人回答环境问题。
- **Python 3.6 是硬门槛**（README Prerequisites 原文 `Python 3.6`）→ 与上面的 2019 pin 互相印证，是**2019 年的运行环境**，在 Ubuntu 24.04 + CUDA 13.2 上无法照原样复现。（详细 pin→sm_120 推导见第 5 节。）
- **权重与检测器都要手动下载且链接易腐**：人脸检测模型要求手动放到 `face_detection/detection/sfd/s3fd.pth`，来源是 `adrianbulat.com`（README 自己提供了备用 SharePoint 链接，说明主链常挂）；模型权重走 Google Drive。
- **许可严格限制商用**：README 原文 **"This repository can only be used for personal/research/non-commercial purposes."** 以及 **"As the models are trained on the LRS2 dataset, any form of commercial use is strictly prohibited."** → **完全不能用商业授权的数字人产品**。

---

## 3. bytedance/LatentSync

### 3.1 2026 状态与版本

| 项目 | 数值 | 证据 |
|---|---|---|
| Stars / Forks | **6090 / 981** | [ungh.cc/repos/bytedance/LatentSync](https://ungh.cc/repos/bytedance/LatentSync)；GitHub Issues 页 aria-label 亦为 "6090 users starred this repository" |
| 创建 / **最后 push** | 2024-12-11 / **2025-06-20** | 同上 → **2026 年零代码更新（截至 2026-09-22 已停更约 15 个月）** |
| GitHub Releases | **空（`{"releases":[]}`）** | 同上。权重走 HuggingFace，不打 GitHub Release |
| 最新版本 | **1.6**（`2025/06/11`） | README Updates：**"`2025/06/11`: We released LatentSync 1.6, which is trained on 512×512 resolution videos to mitigate the blurriness problem."**；HF 仓 [ByteDance/LatentSync-1.6](https://huggingface.co/ByteDance/LatentSync-1.6) |

**"2026 后继版是否改变要求" 的答案：没有 2026 新版；而 1.6 相对 1.5 把显存要求改坏了（8GB → 18GB）。**
- 1.5（2025/03/14）原文：**"improves performance on Chinese videos and reduces the VRAM requirement of the stage2 training to 20 GB"**
- 1.6（2025/06/11）只改了训练分辨率到 512×512，**后果是推理显存从 8GB 抬到 18GB**（见下）。
- **未查到** LatentSync 1.7 或任何 2026 年新版本/新权重。（HF API 在本机超时，我以 README + HF 模型链接为准）

### 3.2 (c) 显存需求 —— 16GB 单卡的判决书就在这里

**【作者自述，README 原文，最权威】**
> Minimum VRAM for inference:
> - **8 GB** with LatentSync 1.5
> - **18 GB** with LatentSync 1.6

链接：[LatentSync README](https://github.com/bytedance/LatentSync)

→ **16GB 单卡 < 官方 18GB 门槛，官方口径下 1.6 不可行。** 要留在 16GB 内，官方口径只允许退回 **1.5（8GB）**。

**训练档（同一 README，供参考）**：`stage1.yaml` 23GB；`stage2.yaml` 30GB；`stage2_efficient.yaml` **20GB**（作者注："suitable for users with consumer-grade GPUs, such as the RTX 3090"）；`stage1_512.yaml` 30GB；`stage2_512.yaml` **55GB**。（[README](https://github.com/bytedance/LatentSync)）

**fp16/权重体积（第三方）**：ComfyUI 封装仓 README 列出 `latentsync_unet.pt (~5GB)`、`stable_syncnet.pt (~1.6GB)`，并称 **"Reduced VRAM Requirements: Optimized to run on 20GB VRAM (RTX 3090 compatible)"**。（[ShmuelRonen/ComfyUI-LatentSyncWrapper](https://github.com/ShmuelRonen/ComfyUI-LatentSyncWrapper)）⚠️ 该仓为第三方封装，且"20GB / RTX 3090"表述自相矛盾（3090 是 24GB），**仅作量级参考**。

**🟢 决定性第三方实测：RTX 5070 Ti 16GB 真的跑起来了（2026 年 7 月）**
Issue **#365 "TensorRT-RTX optimized inference path and reduced-step speed benchmarks"**，**opened on Jul 3, 2026**，作者 Petrus Vermaak（自述 "This work was directed, tested, and validated by Petrus Vermaak... OpenAI Codex served as the implementation and engineering agent"）。**测试系统：`RTX 5070 Ti 16 GB test system`** —— 同为我们关心的显存档位与 Blackwell 世代。基准夹具原文：

> Fixture: official demo video/audio, **9.68s output, 242 frames, 512 face processing, 20 steps, guidance 1.5.**

| Backend | Wall seconds | Sec/output second | Change vs exact20 | **Peak total VRAM** |
|---|---|---|---|---|
| exact20 PyTorch eager | **397.512** | **41.065** | baseline | **15807 MB** |
| optimized20 TensorRT-RTX exact | **268.160** | **27.702** | **32.54% faster** | **12551 MB** |

（逐字引用表格数字；链接：[issue #365](https://github.com/bytedance/LatentSync/issues/365)）

→ **【第三方实测】16GB 单卡在 512×512 / 20 步下可以跑完 LatentSync 1.6**：裸 PyTorch 峰值 **15807 MB**（已占满 16GB 的 96%，极限）、TensorRT-RTX 优化后 **12551 MB**（留有余量）。**官方说的 18GB 是保守门槛，实测 16GB 能压进去，但 PyTorch 裸跑基本贴着天花板**。

**用户侧旁证（都指向 16GB 不够）**：
- Issue **#335**（中文标题，逐字）**"因為顯存不夠，只有16GB，想詢問是否能使用v1.5的版本？"** → 用户因 16GB 不够，主动问能不能退 1.5。（[issue #335](https://github.com/bytedance/LatentSync/issues/335)）
- Issue **#314** 标题 **"8G显存运行不了1.5"**。（[issue #314](https://github.com/bytedance/LatentSync/issues/314)）
- Issue **#284** 标题 **"v1.6 是否可以在8G vRam 运行？有哪些方向可优化以减少 vRam 占用？"**。（[issue #284](https://github.com/bytedance/LatentSync/issues/284)）
- Issue **#278** 标题 **"4090 24G train stage1 OOM"** → 训练侧 24GB 都会 OOM。（[搜索结果页](https://github.com/bytedance/LatentSync/issues?q=is%3Aissue+OOM)）

### 3.3 (d) 端到端时延 / FPS

- **【作者自述】无延迟/FPS 数字**，只有 512×512 分辨率与 `inference_steps [20-50]`、`guidance_scale [1.0-3.0]` 两个可调项（步数越高越慢）。
- **【第三方实测，可用】** Issue #365 同一夹具（242 帧 / 9.68s 输出）：
  - PyTorch eager 20 步：**397.512 秒**完成 9.68 秒视频 → **≈ 0.61 帧/秒**，即**约 41.1 秒墙钟换 1 秒成片**（比实时慢 **≈41×**）。
  - TensorRT-RTX 20 步：**268.160 秒** → **≈ 0.90 帧/秒**，**约 27.7 秒换 1 秒**（仍比实时慢 **≈28×**）。
  - 更长素材（30.08s 真实素材）降步数档：`optimized20` 779.70s、`optimized12` 609.50s、`optimized8` 502.50s、`optimized4` 384.70s；作者自己警告 **"optimized4 was visually plausible in this test but showed the largest objective drift, so I would not present it as equivalent quality"**（降步数是"近似"模式，质量会漂）。
  - 质量守卫：`Engine cosine vs PyTorch 0.9999949036`、`Full-frame SSIM mean 0.981501`、`SyncNet confidence 8.344`（baseline 8.578）。
- **【第三方实测，1.5 时代对照】** Issue **#137** 原文：**"the processing runs at approximately 4 iterations per second (4.17it/s) with 1%-3% GPU utilization"**，并把 Sample frames 16 → 32 后降为 **"2it/s"**；同时报告 **"GPU Memory Usage: 6GB/ 16GB"**。（[issue #137](https://github.com/bytedance/LatentSync/issues/137)）⚠️ 该 Issue 属 1.5 时代、且 4.17 it/s 是**单步迭代速度**而非成片速度，**不可与 #365 的"秒/秒"直接换算**，但它证明**显存并非 1.5 的瓶颈（仅用 6GB），瓶颈是算力与 GPU 利用率低（1%-3%）**。
- **首帧时延（first-frame latency）：未查到** —— 扩散类逐帧批量推理，社区无人报告"首帧"指标，**此项为不适用/无数据**。

### 3.4 (e) 流式 / 可打断

**不支持，且被 Issue 明确暴露。** 架构是**整段扩散去噪**（20 步），没有 chunk 级流水线；Issue **#326** 标题即为 **"How to run inference on longer video?"**（[issue #326](https://github.com/bytedance/LatentSync/issues/326)）→ 长视频要自己切段处理。**无流式输出、无打断机制。**

### 3.5 (f) 中文支持 / 半身全身

- **中文：有专门优化。**【作者自述】1.5 更新说明逐字：**"improves performance on Chinese videos"**（[README](https://github.com/bytedance/LatentSync)）→ 相比 Wav2Lip 的"语言无关但只训练英文"，LatentSync 是**明确宣称对中文视频有优化**的。
- **半身/全身：不支持，仅面部。** 数据管线原文：**"Affine transform the faces according to the landmarks detected by InsightFace, then resize to 256 × 256."**（1.6 为 512×512）→ 模型只在**仿射对齐后的人脸裁剪块**上工作，**输出仅覆盖面部区域**，身体/姿态由原视频保留、不生成。**无半身/全身驱动能力。**（[README](https://github.com/bytedance/LatentSync)）

### 3.6 (g) 坑：torch/CUDA / xformers / flash-attn / 512×512 显存

**🔴 最硬的坑：官方依赖锁死在 cu121，不含 sm_120 内核。**
`requirements.txt` 逐字（[raw](https://raw.githubusercontent.com/bytedance/LatentSync/main/requirements.txt)）：

```
torch==2.5.1
torchvision==0.20.1
--extra-index-url https://download.pytorch.org/whl/cu121
diffusers==0.32.2
transformers==4.48.0
decord==0.6.0
accelerate==0.26.1
einops==0.7.0
omegaconf==2.3.0
opencv-python==4.9.0.80
mediapipe==0.10.11
python_speech_features==0.6
librosa==0.10.1
scenedetect==0.6.1
ffmpeg-python==0.2.0
imageio==2.31.1
imageio-ffmpeg==0.5.1
lpips==0.1.4
face-alignment==1.4.1
gradio==5.24.0
huggingface-hub==0.30.2
numpy==1.26.4
kornia==0.8.0
insightface==0.7.3
onnxruntime-gpu==1.21.0
DeepCache==0.1.1
```

→ **`torch==2.5.1` + `cu121` 官方 wheel 不含 sm_120（Blackwell）内核**。在 RTX 5060 Ti 上照原样装，会撞上与 Duix #624 同型的 `no kernel image is available for execution on the device`。**必须自行升级到 torch ≥ 2.7 + cu128/cu13x**，而**任何 torch 升级都可能打破下面这批老旧 pin**。这是本项目在 sm_120 上最大的工作量来源。

**✅ 好消息：不需要 xformers，也不需要 flash-attn。**
我逐条核对了上面的 requirements：**没有 `xformers`，也没有 `flash-attn`**。→ **不存在 flash-attn/triton 需要现场编译 sm_120 内核的问题**，这一点比很多扩散类 talking-head 项目（如 MuseTalk / 部分 ComfyUI 节点）友好得多。注意力走 diffusers 默认路径。
（若走 **TensorRT-RTX** 路线，Issue #365 已证明在 16GB Blackwell 上可行且比 PyTorch 省 3256MB 峰值显存、快 32.54%，是本项目在 16GB 卡上**最值得走的优化路径**。）

**其他版本坑（逐条）**：
- **`onnxruntime-gpu==1.21.0`**：需要与 CUDA 版本匹配的 ORT 构建；这类 pin 在 CUDA 13.x 宿主上常需换轮子。
- **`insightface==0.7.3` + `face-alignment==1.4.1`**：人脸检测依赖，**要现场编译**（insightface 需 C++ 工具链），是环境搭建的常见失败点。
- **`numpy==1.26.4` + `gradio==5.24.0` + `mediapipe==0.10.11`**：整体是 2025 年上半年的老矩阵，**与 torch≥2.7 的组合未被作者验证过**（作者 2025-06-20 后就停更了）。
- **`DeepCache==0.1.1`**：推理加速用，可能与新版 diffusers 冲突（diffusers 被 pin 在 0.32.2）。
- **`setup_env.sh`**：README 要求 `source setup_env.sh` 来装依赖+下权重，**但我抓取该文件返回空（未核实其内容）** —— 不排除 404 或抓取失败，**此项标注为未核实**。
- **Issue #270 标题 "This USED to work fine for me, but now I get \"No available kernel. Aborting execution.\""** —— 字面即"没有可用内核"，与内核编译/环境漂移高度相关，**但我未逐条核实其是否由 sm_120 引起**，仅作为风险信号列出。（[issue #270](https://github.com/bytedance/LatentSync/issues/270)）
- **Issue 创建被限制**：Issues 页显示 **"New issue — Issue creation is restricted in this repository"** → 结合 15 个月无代码更新，**上游响应能力基本为零**。

### 3.7 LatentSync 面向 RTX 5060 Ti 16GB 的一句话判决

**能跑，但要打破官方依赖并接受"离线批处理"定位**：官方 18GB 门槛把 1.6 判为不合规；实测 16GB（5070 Ti）在 512×512/20 步下峰值 15807MB 可通过，TensorRT-RTX 降到 12551MB。**代价是速度仅约实时的 1/28~1/41，且必须自己把 torch 从 cu121 迁到 cu128+ 才能吃 sm_120。** 它不是"实时数字人"，是"离线 lip-sync 渲染器"。

---

## 4. 三仓横向对照（面向 RTX 5060 Ti 16GB / sm_120）

| 维度 | Duix.Avatar（原 Duix.Heygem） | Wav2Lip | LatentSync |
|---|---|---|---|
| 2026 状态 | 改名后仍活跃（2026-04 有提交） | **停更**（2025-06-22 最后一次 push，README 已商业导流） | **停更 15 个月**（2025-06-20） |
| Stars | 15547 | 13219 | 6090 |
| 官方推理显存 | 未给数字（整机推荐 RTX 4070，16GB+ 为推荐档） | **未给** | **8GB(1.5) / 18GB(1.6)** |
| 16GB 单卡 | ⚠️ **怕是不行**：三容器常驻，无 16GB 实测 | ✅ 几无悬念可跑（推断） | ⚠️ 官方不合规，**实测 15807MB 压线可跑** |
| 成片速度 | 未查到 FPS（官方定义非实时） | 未查到（Issue 普遍抱怨慢） | **27.7~41.1 秒 / 1 秒成片**（5070Ti 16GB 实测） |
| 流式/打断 | ❌ 提交+轮询 | ❌ 整段批处理 | ❌ 整段扩散 |
| 中文 | ✅ 8 语言含中文 | ⚠️ 语言无关，但只训练英文(LRS2) | ✅ **明确优化中文视频** |
| 半身/全身 | 【推断】继承源视频（face2face） | 只改嘴部，身体不动 | ❌ **仅面部** |
| **sm_120 现状** | 🔴 **坏**（Issue #624 Open 零回复，主线镜像无 CC12.0 内核；仅 5090 专用镜像存在，未验证 5060 Ti） | 🔴 **pin 即不可用**（`torch==1.1.0` 为 2019 年 CUDA 9/10 时代，且无 cp38+ 轮子；须整套重写） | 🟡 需自行把 cu121→cu128+（**无 xformers/flash-attn，省一大坑**） |
| 商用许可 | ⚠️ **自定义许可（非开源）**：DUIX.COM COMMUNITY LICENSE，**MAU > 1000 即须申请商业授权**（README 误写"10 万"，差 100 倍） | ❌ **严格禁止商用**（LRS2 数据集限制） | ✅ **Apache License 2.0**（LICENSE 正文已核实） |
| 部署形态 | Docker 三容器 + 客户端 App | 裸 Python 脚本 | 裸 Python 脚本 |

---

## 5. 依赖 pin 对 sm_120 / Blackwell 的含义（本轮重点补充）

### 5.1 事实基线：要原生支持 sm_120 需要什么

**【官方，PyTorch 官方博客】PyTorch 2.7 发布说明逐字**（[pytorch.org/blog/pytorch-2.7](https://pytorch.org/blog/pytorch-2.7/)，180248 字节，我已抓取正文）：
- **"PyTorch 2.7 introduces support for NVIDIA's new Blackwell GPU architecture and ships pre-built wheels for CUDA 12.8."**
- 该小节标题带明确限定词：**"[Prototype] NVIDIA Blackwell Architecture Support"** → 即 2.7 的 Blackwell 支持在当时是**原型级**，不是完全体。
- **"PyTorch 2.7 includes Triton 3.3, which adds support for the Blackwell architecture with torch.compile compatibility."**
- 官方安装命令逐字：**`pip install torch==2.7.0 --index-url https://download.pytorch.org/whl/cu128`**

→ **基线结论：要原生吃 sm_120，torch 必须 ≥ 2.7 且用 cu128（或更新）轮子；cu121 及更早的官方轮子不含 sm_120 内核。**

**【第三方实测，逐字警告文本 —— 索引标题级证据】** Comfy-Org/ComfyUI Issue **#7127** 的页面标题逐字包含运行期警告：

> UserWarning: NVIDIA GeForce RTX 5070 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation. **The current PyTorch install supports CUDA capabilities sm_50 sm_60 sm_61 sm_70 sm_75 sm_80 sm_86 sm_90.**

→ 这就是 **cu121 轮子的实际 arch 列表：最高到 sm_90，没有 sm_120**。触发设备是 **RTX 5070 Ti（同为 sm_120）**，与我们的 5060 Ti 同代同 cap。
⚠️ **诚实标注**：本轮 `github.com` 的 Issue 正文页多次抓取超时（网络问题），上述引文来自**搜索索引返回的页面标题原文**，我**未读取该 Issue 正文**；`pytorch/pytorch` #166794 同理，其标题逐字为 **"[Bug] RTX 5070 Ti (sm_120) not recognized by PyTorch 2.5.1+cu121"**。两者作为"同型 pin 已被报告失败"的证据成立，但均为**标题级、非正文级**证据。
【第三方实测，同型】HybridRobotics/Berkeley-Humanoid-Lite Issue **#49** 标题逐字：**"RTX 5090 (SM_120) fails with default torch/cu121: no kernel image is available"**。

### 5.2 三条路线的 pin 逐条判定

| 仓库 | pin（逐字） | sm_120 可用性 | 必须做什么 |
|---|---|---|---|
| **LatentSync** | `torch==2.5.1` / `torchvision==0.20.1` / `--extra-index-url .../cu121` / `onnxruntime-gpu==1.21.0` | 🔴 **不可用**：2.5.1 早于 2.7，cu121 arch 表止于 **sm_90** | 升 torch ≥2.7 + cu128，并连带处理 5.4 的脆点 |
| **Wav2Lip** | `torch==1.1.0` / `torchvision==0.3.0` / `numba==0.48` / `numpy==1.17.1` / `librosa==0.7.0` | 🔴 **绝对不可用**：2019 年 CUDA 9/10 产物，且无 cp38+ 轮子 | pin 形同废弃，须整体现代化重写 |
| **Duix.Avatar** | 源码**无 pin**——pin 在 **Docker 镜像内部** | 🔴 **已实测失败**：Issue #624 证实镜像不含 CC 12.0 内核 | **用户无法自行解决**，只能等官方重建镜像或自建镜像 |

**特别注意 LatentSync 与 #166794 的 pin 是同型**：`torch 2.5.1` + `cu121` 正是被反复报告在 **RTX 5070 Ti / 5090（sm_120）** 上失败的那一组。**这不是理论风险，而是已在同代硬件上被复现的失败组合。**

### 5.3 Wav2Lip 的 `torch==1.1.0` 到底有多致命（本轮新查，PyPI 硬证据）

【第三方可核实，PyPI JSON API】`torch==1.1.0` 的**全部 9 个 wheel**（`https://pypi.org/pypi/torch/1.1.0/json`）：
- 上传时间 **2019-04-30**（同批：`torchvision==0.3.0` 2019-05-22、`numpy==1.17.1` 2019-08-27、`numba==0.48` 2020-06-30）
- 文件名与 Python 标签：`cp27` / `cp35` / `cp36` / `cp37` —— **没有任何 cp38+ 轮子，也没有 aarch64**；`torch-1.1.0-cp37-cp37m-manylinux1_x86_64.whl`（**manylinux1 = 2010 年 ABI 基线**）
- `requires_python` 字段为空（2019 年的包普遍不声明）

**三层含义**：
1. **装不上**：Ubuntu 24.04 默认 Python 3.12，torch 1.1.0 无对应 wheel，`pip install -r requirements.txt` **必然失败**；若强行从源码编译，2019 年的 CUDA 代码在 CUDA 13.2 的 nvcc 下几乎必然编译失败。
2. **跑不了**：即便人工凑出环境，1.1.0 的 CUDA 轮子属 CUDA 9.0/10.0 时代，**arch 覆盖远低于 sm_120**，必报 `no kernel image is available for execution on the device`（与 Duix #624 同型错误）。
3. **生态全断**：`numba==0.48`（依赖 llvmlite 0.31）、`numpy==1.17.1`、`librosa==0.7.0` 在 Python 3.12 上均无可用轮子。

→ **最终判定：Wav2Lip 的 `requirements.txt` 在 2026 年的 Ubuntu + sm_120 上属于"不可复现文档"。它无法按官方方式部署，只能把 `wav2lip.pth` 权重 + 推理逻辑搬到现代 torch 上重写**（社区各种 fork 与 ComfyUI 节点本质都是这么做的）。这与 README 自述的 `Python 3.6` 前提互相印证，也解释了为何其 Issue 追踪器自 2024 年初起就无人处理环境类问题。

### 5.4 LatentSync 换 torch 时需连带处理的 5 个脆点

**唯一的硬门槛是 torch 2.5.1 → ≥2.7(+cu128)**，但换版本号会牵动一串 2025 年上半年的旧 pin：

1. **cuDNN 大版本：不冲突（好消息）**。ORT 文档逐字：**"PyTorch 2.3 uses cuDNN 8.x, while PyTorch 2.4 or later uses cuDNN 9.x"** → 2.5.1 与 2.7 同为 **cuDNN 9.x**，这一项不需要额外折腾。
2. **`decord==0.6.0`**：视频解码轮子，长期只发布到 cp310/cp311 附近，**新 Python 上常无轮子** → 这类项目的经典断点。
3. **`insightface==0.7.3`**：需现场 C++ 编译（人脸检测依赖），是环境搭建的常见失败点。
4. **`DeepCache==0.1.1` + `diffusers==0.32.2`**：DeepCache 是推理加速补丁，**与新版 diffusers 的组合从未被作者验证**（作者 2025-06-20 后停更）。
5. **`onnxruntime-gpu==1.21.0`**：与你的 **CUDA 13.2 宿主**最容易打架，单独见 5.5。

**✅ 一个重要的好消息**：LatentSync 的 `requirements.txt` **没有 `xformers`，也没有 `flash-attn`** → **不需要为 sm_120 现场编译 flash-attn / triton 内核**。中文资料里常见的那类报错（如 cnblogs 记录的 **"nvcc fatal : Unsupported gpu architecture 'compute_120'"**）**在 LatentSync 的原生路径上不会出现**，只在你想额外加装 flash-attn 时才会撞上。若走 Triton 路线，PyTorch 2.7 已内置 **Triton 3.3**，官方称其 "adds support for the Blackwell architecture with torch.compile compatibility"。

### 5.5 `onnxruntime-gpu==1.21.0` 与 CUDA 13.2 宿主的冲突（本轮新查）

**【官方，ONNX Runtime CUDA EP 文档逐字】**（[onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider](https://onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider)，正文已抓取）：
- **"ONNX Runtime built with CUDA 12.8 are compatible with any CUDA 12.x version; ONNX Runtime built with CUDA 13.0 are compatible with any CUDA 13.x version."**
- 版本表逐字：**"1.26.x-1.21.x | CUDA 12.8 | cuDNN 9.x | Available in PyPI and NuGet. Default GPU package build before 1.27."**
- **"Starting with version 1.27, GPU packages published to PyPI (onnxruntime-gpu) ... are built with CUDA 13.0 by default."**
- 另一条硬约束：**"ONNX Runtime built with cuDNN 8.x is not compatible with cuDNN 9.x, and vice versa."**

→ **`onnxruntime-gpu==1.21.0` 是 CUDA 12.8 + cuDNN 9.x 构建**，而你的宿主是 **CUDA 13.2（跨大版本）**。按 NVIDIA Minor Version Compatibility，12.8 构建只保证与 **12.x** 兼容。
→ 【推断，风险项】**两条路**：① 额外提供 CUDA 12 运行库（如 pip 的 `nvidia-cuda-runtime-cu12` / `nvidia-cudnn-cu12`），否则可能缺 `libcudart.so.12`；② 把 ORT 升到 **≥1.27**（cu13.0 构建）以匹配宿主 —— 但**后者已脱离 LatentSync 的 pin，必须重测**。
→ Blackwell 专门适配：搜索索引显示 microsoft/onnxruntime 有一个 PR **#23928 "Extend CMAKE_CUDA_FLAGS with all Blackwell compute capacity"**（标题逐字），说明 ORT 曾专门为 Blackwell 补齐 CUDA arch 编译标志。⚠️ **该 PR 进入哪个发布版本我未能核实**（PR 页本轮网络不可达），**标注为未核实**；实务上 cu12.8 构建通常可借 PTX JIT 在 sm_120 上跑通，但**需实测确认**。

### 5.6 汇总：三条路线在 RTX 5060 Ti 16GB 上的 sm_120 工作量

| 仓库 | 16GB 上能否跑 | 改动量 | 主要拦路虎 |
|---|---|---|---|
| **LatentSync** | ✅ 实测可行（峰值 15807MB 压线 / TensorRT 后 12551MB） | **中**：换 torch≥2.7+cu128，处理 decord / ORT / DeepCache | 无 flash-attn 坑是最大优势；ORT 与 CUDA 13.2 需调和 |
| **Duix.Avatar** | ❌ 目前不可行（#624 已实测同型失败） | **无法自行解决**（内核编译在官方镜像里） | 官方 #624 零响应；仅 5090 专用镜像，未验证 5060 Ti |
| **Wav2Lip** | ⚠️ 理论上限低但**无法按官方方式部署** | **大**：2019 依赖整套作废，须重写推理脚本 | `torch==1.1.0` 无 cp38+ 轮子；且**严格禁止商用** |

**一句话**：三个项目的 pin 无一为 sm_120 准备。**LatentSync 只差一个 torch 升级（且没有 flash-attn 这个天坑），是三者中唯一在 16GB Blackwell 上被实测跑通的**；Duix 卡在官方镜像、用户无力自救；Wav2Lip 的依赖本身就是 2019 年化石。

---

## 来源清单

**Duix 系**
1. https://github.com/duixcom/Duix-Avatar — 主仓（原 Duix.Heygem），README 硬件/语言/许可/API 口径
2. https://github.com/duixcom/Duix-Avatar/releases — v1.0.5 / v1.0.6 改名说明
3. https://github.com/duixcom/Duix-Avatar/issues/624 — **RTX 5070 Blackwell `no kernel image is available`（2026-09-04 开，Open 零回复）**
4. https://github.com/duixcom/Duix-Avatar/issues/325 — "6G VRAM can't run with lite version"
5. https://github.com/duixcom/Duix-Avatar/issues/251 — 4×24G 显卡推理慢
6. https://raw.githubusercontent.com/duixcom/Duix-Avatar/main/deploy/docker-compose-linux.yml — 三服务定义
7. https://raw.githubusercontent.com/duixcom/Duix-Avatar/main/deploy/docker-compose-5090.yml — 50 系专用镜像
8. https://deepwiki.com/duixcom/Duix-Avatar/6.3-hardware-and-performance-tuning — 硬件/显存表（AI 生成 Wiki，含 5090 = CC 9.0 错误，仅供参照）
9. https://github.com/duixcom/Duix-Mobile — 移动端实时 SDK（120ms / <1.5s / 流式 barge-in）
10. https://github.com/duixcom/duix-skills — 2026-07-14 新建的云端 API skill（**非本地后继者**）
11. https://ungh.cc/repos/duixcom/Duix-Avatar — stars/日期（2026-09-22 读取）
12. https://ungh.cc/repos/duixcom/Duix-Mobile 、 https://ungh.cc/repos/duixcom/duix-skills — 同上

**Wav2Lip**
13. https://github.com/Rudrabha/Wav2Lip — README（商业导流、语言、许可、Python 3.6）
14. https://github.com/Rudrabha/Wav2Lip/issues/584 — 10s/1min 耗时提问（2023-11-10，Open 无答）
15. https://github.com/Rudrabha/Wav2Lip/issues?q=is%3Aissue+VRAM — Issue 追踪器活动停滞证据
16. https://ungh.cc/repos/Rudrabha/Wav2Lip — stars / 最后 push 2025-06-22

**LatentSync**
17. https://github.com/bytedance/LatentSync — README（**1.5=8GB / 1.6=18GB**、训练档、中文优化、面部仿射管线、版本更新日志）
18. https://raw.githubusercontent.com/bytedance/LatentSync/main/requirements.txt — **torch==2.5.1 + cu121**（无 xformers/flash-attn）
19. https://github.com/bytedance/LatentSync/issues/365 — **RTX 5070 Ti 16GB 实测：397.512s/15807MB、268.160s/12551MB（2026-07-03）**
20. https://github.com/bytedance/LatentSync/issues/335 — "因為顯存不夠，只有16GB"
21. https://github.com/bytedance/LatentSync/issues/314 — "8G显存运行不了1.5"
22. https://github.com/bytedance/LatentSync/issues/284 — "v1.6 是否可以在8G vRam 运行？"
23. https://github.com/bytedance/LatentSync/issues/137 — 4.17it/s、GPU 利用率 1%-3%、6GB/16GB
24. https://github.com/bytedance/LatentSync/issues/270 — "No available kernel. Aborting execution."
25. https://github.com/bytedance/LatentSync/issues/326 — 长视频推理
26. https://huggingface.co/ByteDance/LatentSync-1.6 — 1.6 权重（最新版本）
27. https://github.com/ShmuelRonen/ComfyUI-LatentSyncWrapper — 第三方封装（~5GB unet / ~1.6GB syncnet / "20GB VRAM" 表述）
28. https://ungh.cc/repos/bytedance/LatentSync — stars / 最后 push 2025-06-20

**依赖 pin / sm_120 / 许可 专项（本轮补查）**
29. https://raw.githubusercontent.com/Rudrabha/Wav2Lip/master/requirements.txt — **逐字 `torch==1.1.0` / `torchvision==0.3.0` / `numba==0.48` / `numpy==1.17.1` / `librosa==0.7.0`**
30. https://pypi.org/pypi/torch/1.1.0/json — **torch 1.1.0 全部 9 个 wheel，上传 2019-04-30，仅 cp27/cp35/cp36/cp37**
31. https://pypi.org/pypi/torchvision/0.3.0/json 、 https://pypi.org/pypi/numba/0.48/json 、 https://pypi.org/pypi/numpy/1.17.1/json — 同批 2019–2020 上传时间
32. https://raw.githubusercontent.com/duixcom/Duix-Avatar/main/LICENSE — **DUIX.COM COMMUNITY LICENSE AGREEMENT 正文；第 2 条 MAU > 1 thousand 门槛；第 1.b.i / 5.c / 5.b 条附加义务**
33. https://raw.githubusercontent.com/bytedance/LatentSync/main/LICENSE — **Apache License Version 2.0 正文（已核实）**
34. https://pytorch.org/blog/pytorch-2.7/ — **"PyTorch 2.7 introduces support for NVIDIA's new Blackwell GPU architecture and ships pre-built wheels for CUDA 12.8"；"[Prototype] NVIDIA Blackwell Architecture Support"；Triton 3.3**
35. https://github.com/Comfy-Org/ComfyUI/issues/7127 — **UserWarning 逐字含 cu121 的 arch 列表 `sm_50 ... sm_90`（无 sm_120）**（索引标题级证据）
36. https://github.com/pytorch/pytorch/issues/166794 — **"[Bug] RTX 5070 Ti (sm_120) not recognized by PyTorch 2.5.1+cu121"**（索引标题级证据，与 LatentSync pin 同型）
37. https://github.com/HybridRobotics/Berkeley-Humanoid-Lite/issues/49 — **"RTX 5090 (SM_120) fails with default torch/cu121: no kernel image is available"**（索引标题级证据）
38. https://onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider — **ORT 1.21.x–1.26.x = CUDA 12.8 + cuDNN 9.x；1.27 起默认 CUDA 13.0；12.8 构建仅保证与 CUDA 12.x 兼容**
39. https://github.com/microsoft/onnxruntime/pull/23928 — "Extend CMAKE_CUDA_FLAGS with all Blackwell compute capacity"（索引标题级证据；**所入版本未核实**）

**未查到 / 未能核实的项（诚实标注）**
- Duix.Avatar 的 fp16 权重体积、具体推理 FPS/首帧时延、16GB 单卡实测 —— **未查到**
- Duix Docker Hub 镜像体积（`hub.docker.com` 本机多次超时） —— **未查到**
- Wav2Lip 的显存/耗时具体数字 —— **未查到**（仅有投诉、无数字）
- `duixcom` org 完整仓库列表 —— 官方 API 限流（core remaining=0），改用 ungh 逐仓探测，**可能遗漏 org 下其他未探测的仓**
- LatentSync 是否存在 1.7 —— **未查到**（README 停在 1.6，HF API 超时）
- **精确到"最后一次 commit 时间"** —— `commits/<branch>.atom` 方案本轮多次超时失败，故全文采用 ungh.cc 的 **`pushedAt`**（Duix-Avatar `2026-04-21T07:06:36Z`、Wav2Lip `2025-06-22T02:41:21Z`、LatentSync `2025-06-20T07:36:58Z`），**该字段为"最后一次 push"而非严格最后一次 commit**，量级与结论不受影响
- 上述 #7127 / #166794 / #49 / PR #23928 的**正文内容** —— 本轮 `github.com` HTML 抓取持续超时，仅有**页面标题逐字**，已逐条标注为"索引标题级证据"
- Duix README 与 LICENSE 的口径差异（100,000 vs 1,000 MAU）—— 已确认存在，但**官方未解释哪个为准**；本报告按 **LICENSE 正文**为准（法律文本优先）
- `hub.docker.com` 镜像 tag 体积与更新日期 —— **未查到**（多次超时）
