# 集群 C：实时数字人（Talking Head / Avatar）开源项目 × RTX 5060 Ti 16GB 选型调研

- **核实日期**：2026-09-22（所有 stars / 最后提交均为当日抓取）
- **目标硬件**：RTX 5060 Ti **16GB**、**sm_120（Blackwell）**、32GB RAM、Ubuntu、驱动 610.43.02、CUDA 13.2
- **数据来源约定**：
  - `[官方自述]` = 项目 README / 官方文档原文
  - `[作者实测]` = 项目作者本人给出的实测
  - `[第三方实测]` = 用户 issue / 博客 / 基准页的实测
  - `[推算]` = 由可验证数字推导，非直接引用
  - **`未查到` = 无可验证来源，本报告不编造**
- **统计口径**：stars / license / last-commit **逐字引用 shields.io 返回值**（见各节 (a)），最后提交的精确时间戳另用 GitHub commits Atom feed 核对（`https://github.com/<owner>/<repo>/commits/<branch>.atom`）。
- **本次抓取的环境坑（方法论局限）**：GitHub REST API 未认证额度 60/hr，本次会话中途即 `rate limit exceeded`，故部分 issue/PR 正文改用 github.com 页面内嵌 JSON 或镜像解析；star 数改用 shields.io；个别页面被反爬（CSDN 两篇只返回 2KB 壳页）。

---

## 0. 五条跨项目结论（先看这个）

1. **Blackwell / sm_120 的硬门槛是 PyTorch ≥ 2.7 + cu128。** [PyTorch 2.7 官方发布说明](https://pytorch.org/blog/pytorch-2.7/)逐字：「PyTorch 2.7 introduces support for **NVIDIA's new Blackwell GPU architecture** and ships **pre-built wheels for CUDA 12.8**」「PyTorch 2.7 includes **Triton 3.3**, which adds support for the Blackwell architecture with torch.compile compatibility」；安装命令逐字 `pip install torch==2.7.0 --index-url https://download.pytorch.org/whl/cu128`。
   ⚠️ **重要限定词**：该小节的官方标题是「**`[Prototype]` NVIDIA Blackwell Architecture Support**」——即 2.7 的 Blackwell 支持当时标注为**原型级**，非完全成熟。
   `[第三方·索引标题级证据]` **cu121 轮子的实际 arch 列表止于 sm_90**：ComfyUI issue #7127 页面标题逐字含运行期警告「The current PyTorch install supports CUDA capabilities **sm_50 sm_60 sm_61 sm_70 sm_75 sm_80 sm_86 sm_90**」——**没有 sm_120**，触发设备正是 **RTX 5070 Ti**。同型失败另有 pytorch/pytorch issue **#166794** 标题逐字「[Bug] **RTX 5070 Ti (sm_120) not recognized by PyTorch 2.5.1+cu121**」（**与 LatentSync 的 pin 完全同型**）与 Berkeley-Humanoid-Lite issue #49 标题逐字「RTX 5090 (SM_120) fails with default torch/cu121: no kernel image is available」。（**诚实标注：本轮 issue 正文页持续超时，以上为搜索索引返回的页面标题逐字，未读正文。**）
   → **凡把 torch 钉在 <2.7 或 cu121/cu117/cu124 的项目（Ultralight、LatentSync、Wav2Lip），原样在 5060 Ti 上都会 `no kernel image is available`。**
2. **ONNX Runtime 是第二个深坑。** `[第三方]` [Natfii/onnxruntime-gpu-blackwell](https://github.com/Natfii/onnxruntime-gpu-blackwell) 原文：「The official PyPI `onnxruntime-gpu` package **does not include `sm_120` kernels**, so `CUDAExecutionProvider` is **unavailable on Blackwell cards and all operations fall back to CPU**.」另有 [microsoft/onnxruntime issue #27621](https://github.com/microsoft/onnxruntime/issues/27621) 标题即「[CUDA EP] **Silent deadlock** … on **Blackwell GPU (RTX 5060, sm_120, Windows)**」。官方[版本表](https://onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider.html)逐字：「Starting with version **1.27**, GPU packages published to PyPI … are **built with CUDA 13.0** by default. Older GPU package versions are built with **CUDA 12.8** by default.」→ **依赖 onnxruntime 的项目（LivePortrait、LatentSync、Ultralight 流式）必须换到 ≥1.27 或第三方重建 wheel。**
3. **⚠️「官方 README 说支持 50 系」≠「实际能用」——Duix.Avatar 就是反例。** README 明确写了 50 系部署方案，但 2026-09-04 的 [issue #624](https://github.com/duixcom/Duix-Avatar/issues/624) 报的正是 Blackwell `no kernel image is available`，**至今 Open 且零回复**。**选型必须以 issue/实测为准，不能只看 README 声明。**
4. **维护活跃度是被低估的选型维度。** sm_120 新平台的坑只能靠上游修。本批中 **MuseTalk（最后提交 2025-09-26）、LatentSync（2025-06-20）、Wav2Lip（2025-06-22）均已停更近一年**，其 2026 年的 Blackwell 适配案例**全部来自社区自发、未合并进主干**——选它们就得自己维护补丁分支。
5. **「实时」这个词在本批项目中水分很大。** 只有 **LiveTalking** 与 **OpenAvatarChat** 是真正的流式双工框架；MuseTalk 项目成员本人确认**不支持音频流式输入**；Duix.Avatar 官方 README 自述是「**non-real-time video synthesis**」；LatentSync / Wav2Lip / LivePortrait 都是离线批处理。

---

## 1. lipku/LiveTalking

### (a) 统计口径（shields.io 原始值，2026-09-22 抓取）
| 项 | 原始值 |
|---|---|
| stars | `{"label":"stars","message":"9.6k",...}` → **9.6k** |
| license | `{"label":"license","message":"Apache-2.0","color":"green",...}` → **Apache-2.0**（`LICENSE` 实取到 Apache License 正文，确认） |
| last commit | `{"label":"last commit","message":"september",...}` → 逐字 **`september`**；Atom 精确值 **2026-09-13T11:56:50Z**（距调研日 9 天） |

### (b) 实现类型
**2D 视频驱动 + 实时流式推流框架**。不是单一模型，而是 `wav2lip` / `musetalk` / `Ultralight-Digital-Human`（+ `ernerf`）三种后端的**流式封装**，加 WebRTC/RTMP/虚拟摄像头推流管线与 LLM+TTS 对话链路。[官方 README](https://raw.githubusercontent.com/lipku/LiveTalking/main/README.md)「Features」逐字：「支持多种数字人模型: ernerf、musetalk、wav2lip、Ultralight-Digital-Human」。

### (c) 显存需求 + 16GB 单卡能否跑
**① [作者自述·估算] 三后端显存**（作者 lipku 本人的技术博客「恒中」，[拆解 LiveTalking 的资源消耗](https://www.cnblogs.com/livetalking/articles/22134438)，表格标题明确标注「显存占用（**估算**）」）：

| 模型 | 技术路线 | 显存占用（估算） | 精度 | 输入分辨率 |
|---|---|---|---|---|
| Wav2Lip | CNN 编解码 + 空间注意力 | **~200MB** | FP32 | 256x256 |
| MuseTalk | UNet (Diffusion) + VAE + Whisper | **~1.2GB（FP32）/ ~600MB（FP16）** | FP16 | 256×256 |
| UltraLight | 轻量 UNet + HuBERT | **~1.5GB** | FP32 | 160×160 |

⚠️ 该表是**读源码的估算值**，不是实测峰值。**实测峰值看下面两条。**

**② [作者实测] 三模型 4090 实测显存与吞吐**（作者 lipku 的 CSDN 文章 [livetalking实时数字人多并发](https://blog.csdn.net/lipku/article/details/148593974)，逐字）：
> 「以下数据都是在**4090 24G显卡**上测试得出。」
> 「wav2lip模型推理速度很快…**占用显存在1.3G，推理速度能达到750fps**。按照每路视频25fps计算，在同时说话时能支持30路。」
> 「musetalk对显存和gpu计算性能都要求很高，**占用显存12G，推理速度60fps**。能够支持2路同时说话。」
> 「ernerf对显存占用不太高，**大概2G左右**…推理速度45fps。」

**③ 🔴 [第三方实测] MuseTalk 后端的 no_grad 显存巨坑 —— 17.62GB → 3.9GB**
[PR #612](https://github.com/lipku/LiveTalking/pull/612) 标题逐字：「fix: MuseReal.inference_batch 缺 `@torch.no_grad()`，实时推理显存 **17.6G→3.9G**」，**已于 2026-08-20 被作者合并**。正文实测逐字：「实测（**RTX 5090**，musetalk，**batch_size=16**）：torch allocated 峰值 **17.62 GB**」。
未修复时原始报错见 [issue #192](https://github.com/lipku/LiveTalking/issues/192) 逐字：「`RuntimeError: CUDA out of memory. Tried to allocate 256.00 MiB (GPU 0; 23.65 GiB total capacity; 17.77 GiB already allocated; 45.19 MiB free; 18.47 GiB reserved in total by PyTorch)`」；社区解法逐字：「在 musereal.py 的 inference 方法上添加 `@torch.no_grad()` 注解；启动app.py时，设置 **batch_size为8** 或更小，我设置8可以正常运行」。

**→ 16GB 单卡结论**：
- `--model wav2lip`：**毫无压力**（实测 1.3GB）。
- `--model musetalk`：**能跑，但有硬前提**——所用版本**必须包含 2026-08-20 合并的 PR #612**，否则默认 `batch_size=16` 下峰值 17.6–18.5GB，**16GB 卡必然 OOM**。修复后峰值 3.9GB，余量充足。用旧版则手动加 `@torch.no_grad()` 并把 `batch_size` 降到 8。
- `Ultralight`：~1.5GB（估算），无压力。

### (d) 端到端延迟：首帧 + FPS
**[作者自述] 官方 FPS 表**（README「6. 性能指标」逐字）：

| 模型 | 显卡 | FPS |
|:------|:------|:----|
| wav2lip256 | RTX 3060 | **60** |
| wav2lip256 | RTX 3080Ti | **120** |
| musetalk | RTX 3080Ti | **42** |
| musetalk | RTX 3090 | **45** |
| musetalk | RTX 4090 | **72** |

官方口径逐字：「后端日志 `inferfps` = GPU 推理帧率, `finalfps` = 最终推流帧率，**两者均需 >=25 才算实时**」；「wav2lip256 推荐 RTX 3060 及以上；musetalk 推荐 RTX 3080Ti 及以上」。张量形状逐字（同上博客）：`audiofeat_batch` shape `(batch_size, 1, 80, 16)`、`img_batch` shape `(batch_size, 6, 192, 192)`，**每帧 40ms 预算（25fps）**。
⚠️ 官方 FPS 表**未标注分辨率**，且**无任何 50 系/sm_120 数据**。

**首帧延迟（first-frame latency）：未查到**官方数字。最接近的第三方记录是 [issue #368](https://github.com/lipku/LiveTalking/issues/368) 逐字：「webrtc从点击连接到前端有画面显示的过程中，有时候超级慢要几分钟，有时候只要6s的样子」——**这是 WebRTC 链路建立耗时，不是模型首帧推理耗时，不可当模型延迟引用**。

### (e) 流式 / 可打断
**✅ 都是核心卖点。** [官方 README](https://raw.githubusercontent.com/lipku/LiveTalking/main/README.md)「Features」逐字：「**3. 支持数字人说话被打断**」「5. 支持 WebRTC、RTMP、虚拟摄像头输出」「7. 支持多并发」「6. 支持动作编排：不说话时播放自定义视频」。架构上 `/human`（文本，echo/chat 模式）与 `/humanaudio`（音频文件）接口，每连接分配唯一 `sessionid` 支持多用户并发。
（注：**支持打断，但未见"双工打断"声明**；对比 OpenAvatarChat 0.6.0 明确支持手动+双工打断。）

### (f) 中文 / 半身-全身
- **中文：完整支持**。中文为一等公民文档语言；TTS 模块化支持 EdgeTTS、GPT-SoVITS、CosyVoice、腾讯云等中文方案；LLM 对接 Qwen。2026 年仍有中文 issue 活跃（[issue #618](https://github.com/lipku/LiveTalking/issues/618) 请求接入 audio.cpp 做本地 ASR/TTS）。
- **全身：明确支持。** Features 逐字：「**4. 支持全身视频拼接**」。半身通过 `avatar.html` 上传视频自动生成形象实现（README 未单列「半身」词条）。

### (g) 已知坑
- ✅ **官方验证环境恰好是 sm_120 友好档**：README「1. 安装」逐字：「已在 **Ubuntu 22.04、Python 3.12、PyTorch 2.9.1、CUDA 12.8** 测试通过」，安装命令 `pip install torch==2.9.1 ... --index-url https://download.pytorch.org/whl/cu128`。**cu128 + torch 2.9.1 含 sm_120 内核**，与目标机 CUDA 13.2（向下兼容）兼容。
- **依赖极轻，无 flash-attn / triton / onnx / tensorrt 依赖** → **不存在 flash-attn/triton 编译坑，也无 ONNX/TensorRT 版本坑**。[requirements.txt](https://raw.githubusercontent.com/lipku/LiveTalking/main/requirements.txt) 中 torch 相关行为注释状态，torch 由用户按 CUDA 版本自行安装。
- **sm_120 专项 issue：未查到**（`repo:lipku/LiveTalking sm_120` 搜索 total 0，`Blackwell` 亦无结果）→ 属「**无已知 issue**」，**不等于"已明确适配"**。
- **必须注意版本**：musetalk 路径的 no_grad 修复直到 **2026-08-20** 才合并（PR #612）。这是本仓库最大的历史坑。
- **CPU 与 RAM 才是并发瓶颈**（[作者博客](https://www.cnblogs.com/livetalking/articles/22134438)）：
  - WebRTC 默认走 `libx264` **纯软件编码**（非 NVENC），1080p 单路持续占用约 **0.45 个 x86 核心**，多路线性叠加；5 路默认 `max_session` 约耗 **2.0–2.5 核**。反压机制逐字：`if buffer_size >= 5: time.sleep(0.04*buffer_size*0.8)`。
  - **avatar 帧缓存吃 RAM**：10 秒 1080p 素材 ≈ **1.61GB RAM**（逐字算法：250 帧 × 6.2MB ≈ 1.56GB + 人脸缓存 49MB）。**默认 `max_session = 5` 路**。→ 32GB RAM 机器多路时**内存比显存更紧张**。
  - NVENC 硬编并发上限表（3060=3 路、3080Ti/3090=3 路、4090=5-8 路、A5000/A6000=20+ 路、**A100/H100=0 路（无 NVENC 编码器）**）；**表中无 50 系**，`[推算]` 5060 Ti 为 1 个 NVENC / 3 路，**未验证**。
- 官方 FAQ 独立站：[doc.livetalking.ai/docs/faq](https://doc.livetalking.ai/docs/faq/)。

---

## 2. duixcom/Duix.Heygem → **已由官方改名为 duixcom/Duix-Avatar**（🔴 但 sm_120 当前是坏的）

> 🔴 **2026 状态变更（双重）**：
> **(1) 改名**：`https://github.com/duixcom/Duix.Heygem` 返回 **301**，`curl -w "%{url_effective}"` 实测最终 URL = **`https://github.com/duixcom/Duix-Avatar`**。`[第三方核实]` **铁证是 Release v1.0.5（2025-08-15）原文逐字**：「The original project name **"HeyGem" has now been officially changed to "Duix.Avatar"**.」，v1.0.6 补充「It's just a name change.」→ **同一 repo id（907627874）改名，非新仓**。
> **(2) 🔴 sm_120 当前不可用**：见 (g)，[issue #624](https://github.com/duixcom/Duix-Avatar/issues/624) **Open 零回复**。

### (a) 统计口径（shields.io 原始值）
| 项 | 原始值 |
|---|---|
| stars | shields `{"label":"stars","message":"16k",...}` → **16k**；`[第三方]` 精确 **15,547** |
| license | `{"label":"license","message":"not identifiable by github","color":"lightgrey",...}` → **`not identifiable by github`**；实际 `LICENSE` 首行逐字为 **「DUIX.COM COMMUNITY LICENSE AGREEMENT」**（自定义，非 SPDX 标准） |
| last commit | `{"label":"last commit","message":"april",...}` → 逐字 **`april`**；Atom 精确值 **2026-04-21T07:06:08Z** → **到 9 月已 5 个月无更新** |

### (a2) 同组织生态与其他仓库（2026 状态）
从 [`github.com/orgs/duixcom/repositories`](https://github.com/orgs/duixcom/repositories) 抓取 + 逐仓探测：

| 仓库 | stars | 最后提交 | 说明 |
|---|---|---|---|
| [duixcom/Duix-Avatar](https://github.com/duixcom/Duix-Avatar) | 16k / 15,547 | 2026-04-21 | 即原 Duix.Heygem；**官方自述非实时** |
| [duixcom/Duix-Mobile](https://github.com/duixcom/Duix-Mobile)（原 `Duix.mobile`，301 重定向） | **8.3k / 8,252** | **2026-08-05** | **Duix 组织内 2026 最活跃**；移动端实时 SDK（Android/iOS/Pad/车机/VR/IoT）；README 称 **<1.5s 延迟 / 手机端 120ms / 支持流式 barge-in** |
| [duixcom/duix-skills](https://github.com/duixcom/duix-skills) | 4 | 2026-08-12 | ⚠️ **是云端 API skill**（要 `DUIX_API_KEY` + 预付费积分 + 订阅套餐），**不是本地 16GB 部署的后继者** |
| `duixcom/Duix.Heygem.Android` | — | — | **404 不存在** |
| `duixcom/Duix-Reface` | — | — | **404 不存在**（Duix-Mobile README 仍在推荐但已失效） |

→ **2026 继任关系判断**：`Duix.Heygem` → **`Duix-Avatar`（服务端，非实时，5 个月未更新）**；移动端走 **`Duix-Mobile`（2026-08 活跃，支持流式+barge-in）**。**没有出现取代 Heygem 的第三代服务端仓库**；`duix-skills` 是云 API 而非本地部署后继者。

### (b) 实现类型
**2D 视频驱动数字人 + Docker 打包交付。** README 逐字：「`docker-compose up -d` … if you want to use the lite version, execute `docker-compose -f docker-compose-lite.yml up -d`」；镜像 `guiji2025/duix.avatar`、`guiji2025/fish-speech-ziming`、`guiji2025/fun-asr`。
🔴 **官方明确定义自己不是实时的**，README 逐字：「Duix.Avatar's digital human realizes digital human cloning and **non-real-time video synthesis**」→ 实时交互被刻意排除到 duix.com 云端。

### (c) 显存需求 + 16GB 单卡能否跑
**[官方自述] 硬件要求**（README「Hardware Requirements」逐字）：

| 项 | 官方值 |
|---|---|
| CPU | 建议 13th Gen Intel Core **i5-13400F** |
| 内存 | **32GB**（Ubuntu 段特别标注 **「32G or more (necessary)」= 必要**） |
| 显卡 | **RTX 4070**（推荐配置） |
| 硬盘 | C 盘 **>100GB**；D 盘 **>30GB** |
| Docker 模型下载流量 | 「download will consume about **70GB** of traffic」 |

**具体显存 GB：官方未给出，未查到**；**fp16 权重体积：未查到**（权重内嵌 Docker 镜像，仓库不发布权重）。
`[第三方整理·可信度中]` DeepWiki 给：Min RTX 4070 (8GB VRAM) → Recommended RTX 4090/5090 (16GB+ VRAM)；调优表「12-16GB → `max_split_size_mb 512`」。⚠️ **该 Wiki 是 AI 生成，其表把 5090 的 Compute Capability 错标为 9.0（实为 sm_120），需警惕其可信度。**
`[第三方]` [issue #325](https://github.com/duixcom/Duix-Avatar/issues/325) 标题「6G VRAM can't run with lite version」。

**→ 16GB 单卡结论**：⚠️ **未查到任何 16GB 单卡实测报告 → 列为高风险待实测。** 官方推荐卡是 **RTX 4070（12GB）**，容量上 16GB 不低于推荐档；但**三容器（TTS + ASR + 视频合成）常驻**且 TTS 还要做声音克隆训练，且 **32GB RAM 是官方标注的「必要」项**（目标机恰好打平，无余量）。**在 sm_120 上更有 (g) 的阻塞性问题。**

### (d) 端到端延迟：首帧 + FPS
- 🔴 **官方定义就是非实时**（README 逐字「**non-real-time video synthesis**」）→ **具体 FPS / 首帧：未查到**。
- `[第三方]` 反证 [issue #251](https://github.com/duixcom/Duix-Avatar/issues/251)：「我电脑上有4张24G的显卡，**推理速度很慢**」。
- API 形态是**提交 + 轮询**：`/easy/submit`、`/easy/query?code=`。
- ⚠️ **不要误用 Duix-Mobile 的 120ms 数字**——那是**手机 SoC 端侧渲染**，与 Duix-Avatar 服务端无关。

### (e) 流式 / 可打断
**❌ 不支持。** README 自述非实时；TTS 参数逐字 `"streaming": false, // Fixed parameter`；无 barge-in/双工。（**Duix-Mobile 才支持流式 + barge-in**。）

### (f) 中文 / 半身-全身
- **中文：支持。** README 逐字：「Scripts support **eight languages** - English, Japanese, Korean, **Chinese**, French, German, Arabic, and Spanish.」
- **半身/全身**：`[推断]` 技术路线是 **face2face 换脸式驱动**（数据目录 `~/duix_avatar_data/face2face:/code/data`）→ 输出构图继承源视频，**半身/全身取决于源视频**；但**非文档明示**。aigcpanel 对它的描述逐字为「**Heygem** | 全身数字人驱动」。
- 客户端：官方安装包见 [duixcom/Duix.Avatar releases](https://github.com/duixcom/Duix.Avatar/releases)。

### (g) 已知坑 —— 🔴 **sm_120 上当前报错且无人修（本报告最重要的反面教材）**
🔴 **[第三方实测] Issue #624 标题逐字**：「**RTX 5070 (Blackwell) not supported — CUDA kernel error on Docker backend**」，**opened on Sep 4, 2026，状态 Open，页面零回复**。原文逐字：
> 环境：RTX 5070 Laptop / **Compute Capability: 12.0** / Driver 592.15
> `RuntimeError: CUDA error: no kernel image is available for execution on the device`
> 「Root cause: The Docker images are compiled with a **PyTorch version that doesn't include CUDA kernels for compute capability 12.0**. PyTorch 2.7+ added Blackwell support.」

来源：[issue #624](https://github.com/duixcom/Duix-Avatar/issues/624)。
→ **对 RTX 5060 Ti（同 sm_120）的直接含义**：**默认三个 Docker 镜像的主线编译目标不含 sm_120**，容器会直接抛 `no kernel image is available`。**维护者截至 2026-09-22 未回应。**
→ README 另提供独立的 `deploy/docker-compose-5090.yml`（镜像 `guiji2025/duix.avatar-5090` + `fish-speech-5090`），官方声明逐字：「For 50 series graphics cards (**tested and also works for 30/40 series with CUDA 12.8**) Uses the **official preview version of PyTorch**」+「[Nvidia 50 Series GPU Version Notice] **Tested and verified on 5090 GPU**」。→ **走 5090 专用镜像"理论上"是给 Blackwell 的路径，但官方只在 5090 上验证过，5060 Ti 是否可用未查到任何证据；且 #624 证明普通镜像确实没有 sm_120 内核。这是 16GB 选型里最大的未知风险点。**

**其他坑**：
- 🔴 **许可证表述前后冲突（必须注意）**：
  - `LICENSE` 文件第 2 条逐字：「If … the Monthly Active Users … is greater than **1 thousand** … or … greater than **1 thousand Monthly Active Users, you must request a commercial license** from DUIX.COM … and you are **not authorized** to exercise any of the rights under this Agreement unless or until DUIX.COM otherwise expressly grants you such rights.」
  - 但 README 对比表逐字：「Commercial Authorization: **Supports global free commercial use (enterprises with more than 100,000 users or annual revenue exceeding 10 million USD need to sign a commercial license agreement)**」
  - → **两者阈值差 100 倍（1,000 MAU vs 100,000 用户 / $10M 营收）**。官方未说明以哪个为准 → **商用合规建议按 LICENSE 正文（1000 MAU）从严执行**。（`LICENSE` 正文已抓取全文 7120 字节核对。）
- 🔴 **LICENSE 里其他易被忽略的硬性附加义务**（均逐字有据）：
  - **第 1.b.i(B)**：须在网站 / UI / 博客 / 关于页 / 产品文档**显著展示「Built with DUIX.COM」**，并须在 ToS/EULA 中声明产品基于 DUIX.COM 技术。
  - **第 1.b.i**：若用其材料训练/微调并**分发模型**，逐字要求「you shall also include **"DUIX.COM" at the beginning of any such AI model name**」→ **模型名强制前缀**。
  - **第 1.b.iii**：分发须保留 `Notice` 文本文件（署名见下）。
  - **第 5.c**：使用即授予 DUIX **永久/全球/非独占/免版税/不可撤销**许可，可用你的案例做营销，**"without any compensation or attribution to you"**。
  - **第 5.b**：你若起诉 DUIX 侵权，**许可自起诉之日起自动终止**。
- **对照**：**LatentSync 的 LICENSE 正文已核实 = 标准 `Apache License Version 2.0`**（本批三者中唯一真正宽松的开源许可）。
- 必须保留署名：「You must retain … the following attribution notice within a "Notice" text file: "DUIX.COM is licensed under the DUIX.COM Community License, Copyright © DUIX.COM Platforms, Inc. All Rights Reserved."」
- **部署重**：三容器全部 `runtime: nvidia`，其中两个 **`privileged: true`**，`shm_size: '8g'`（`[第三方]` DeepWiki 把 "Bus Error (Core Dumped)" 归因于 shm 不足 → 建议 12g/16g；1080p/60s 建议 8GB、4K 建议 12GB+）；**必须装 NVIDIA Container Toolkit + `nvidia-ctk runtime configure --runtime=docker`**；**无 CPU 回退**；Ubuntu 只在 **22.04 / 内核 6.8.0-52-generic** 完整验证过（**24.04 属未验证**）。
- **官方自评开源版**逐字：「Lip Sync Effect: **Usable effect**」（云端 API 才是 "Stunning"）、「Iteration Speed: **Slow updates, bug fixes depend on the community**」。
- Docker Hub 镜像体积本机多次超时未取到 → **未查到**。

---

## 3. anliyuan/Ultralight-Digital-Human（+ 继任项目 FeatherTalk）

### (a) 统计口径（shields.io 原始值）
| 项 | 原始值 |
|---|---|
| stars | `{"label":"stars","message":"2.6k",...}` → **2.6k** |
| license | shields `{"label":"license","message":"not specified","color":"lightgrey",...}` → **`not specified`**。⚠️ **存在不确定性**：README 徽章写的是 Apache-2.0（`./LICENSE` 链接），但本次 `curl` 该 LICENSE 文件返回 **HTTP 000（连接失败，非 404）** → **既不能断言有、也不能断言无，标为「存疑/未验证」** |
| last commit | `{"label":"last commit","message":"july",...}` → 逐字 **`july`**；Atom 精确值 **2026-07-22T12:29:23Z** |

> ⚠️ **默认分支陷阱**：本仓库默认分支是 **`master`**；**`main` 分支不存在，其 README 返回 404**。抓取请用 `master`。

### (a2) 继任项目 FeatherTalk（2026 新增，重点）
README **首屏第一段**逐字：「## New Project: FeatherTalk」「**c++流式推理代码已开源到 [FeatherTalk](https://github.com/anliyuan/FeatherTalk)**」「It is a **cleaned-up successor** to Ultralight Digital Human, focused on a **lighter audio encoder, easier training, and mobile-friendly deployment**.」

| FeatherTalk 项 | 值 |
|---|---|
| stars | shields 逐字 **`93`** |
| license | **Apache-2.0**（LICENSE 文件与徽章一致） |
| last commit | 逐字 **`september`**；Atom 精确值 **2026-09-04T22:34:19Z** |

README 逐字关键点：「**模型足够小**：当前视觉模型输入为 **144×144**，参数量约 **5.46M**。并且我自己重新训练了一个全新的超轻量级音频编码器」「**推理更高效**：FeatherHuBERT 对整段音频只计算一次…也可以改为流式推理」「**真正可部署**：同时提供 Python 训练/推理和 **C++/MNN** 接入路线，**可纯 CPU 运行**」；Ultralight README 补充逐字：FeatherHuBERT「相比 HuBERT Large，音频侧计算量约降低 **40 倍**」。
⚠️ **FeatherTalk 自身 README 本次抓取返回 0 字节（`main`/`master` 均如此）→ 其显存/延迟数字未查到。**

### (b) 实现类型
**2D 视频驱动的个性化轻量数字人**（每人一段视频训一个专属模型），主打移动端实时。**不是** 3D 参数化、**不是** 3DGS。

### (c) 显存需求 + 16GB 单卡能否跑
**本仓库实测显存：未查到**（README/官方文档均无显存数字，也未找到第三方实测）。可用旁证：
- `[官方自述]` 模型体积极小：作者称新版「**audio + UNet 两个模型的总体积控制在 1M 以内**」；FeatherTalk 视觉模型参数 **5.46M**（144×144 输入）。
- `[他人估算]` LiveTalking 作者把 UltraLight 显存估为 **~1.5GB（FP32）**、输入 160×160（[来源](https://www.cnblogs.com/livetalking/articles/22134438)）。**这是 LiveTalking 作者的估算，非本仓库自述。**
- `[第三方实测]` **真正的瓶颈是内存（RAM）不是显存** —— [issue #120](https://github.com/anliyuan/Ultralight-Digital-Human/issues/120) 逐字：「有没有发现内存占用巨高」「1分半的1080p视频训练完推理，**内存占用近10G**，虽然gpu算力要明显下降可是这内存占用显得很不ultra light啊」。→ 32GB RAM 单路可用，**并发需谨慎**。
- `[官方自述]` 训练默认 `--batchsize 1`、`--epochs 200` → 训练端显存需求极低（**具体数字未查到**）。

**→ 16GB 单卡结论**：**显存侧必然够用**（移动端级模型），但**无可引用的实测显存数字，本报告标注为「未查到实测值，推断极小」**。真正的限制在**软件栈兼容性**（见 (g)）与**内存占用**。

### (d) 端到端延迟：首帧 + FPS
- `[作者实测]` **每帧 <10ms**（README「关于流式推理」逐字）：「实测在**2080**这样的机器上**多个并发时每帧音频处理+视频处理耗时10ms以内**，需要将模型转为onnx」→ 10ms/帧 < 40ms（25fps）预算，**实时性有余量**；**前提是必须转 ONNX**。
- `[作者自述]` 编码器取舍：「hubert performs better, but **wenet is faster and can run in real time on mobile devices**」。
- **首帧延迟：未查到。推理 FPS 具体数值：未查到**（只有「10ms/帧」这一时间口径）。

### (e) 流式 / 可打断
- `[官方自述]` **支持流式，但代码不完善**。README 逐字：「这个模型是支持流式推理的，但是**代码还没有完善**」。实践中社区仍在催（[issue #103](https://github.com/anliyuan/Ultralight-Digital-Human/issues/103) 标题逐字：「还没有开源实时流吗？在不开源我要开源我的自己写的接的实时流的了」）。
- **继任项目 FeatherTalk 已开源 C++ 流式推理**（见 (a2)）。
- **可打断（barge-in）：未查到。** 本仓库无打断机制；该能力由 LiveTalking 那类上游框架提供。

### (f) 中文 / 半身-全身
- **中文：完整支持**（天然）。README 主文档为中文；官方 demo 用中文口播；链路走 wenet/hubert 中文友好特征。
- **半身 / 全身：未查到明确支持。** README 训练数据要求逐字：「必须保证视频中**每一帧都有整张脸露出来**的人物」→ 属**面部/口播半身**场景；开源推理脚本无全身拼接能力（对比 LiveTalking 明确有「全身视频拼接」）。
- 效果警告逐字：「如果你视频中声音质量比较差的话，效果大概率不会好…建议录制视频时候使用外接麦克风」。

### (g) 已知坑
- 🔴 **sm_120/Blackwell：高风险，官方环境原样不可用。** README 指定环境逐字：
  ```bash
  conda create -n dh python=3.10
  conda install pytorch==1.13.1 torchvision==0.14.1 torchaudio==0.13.1 pytorch-cuda=11.7 -c pytorch -c nvidia
  ```
  **PyTorch 1.13.1 + CUDA 11.7 不含 sm_120 内核**，在 RTX 5060 Ti 上会直接 `no kernel image is available`。作者仅逐字说「我是在1.13.1版本的pytorch跑的，**其他版本的pytorch应该也可以**」——**是猜测，非验证结论**，必须自行迁移到 cu128 栈并重测。
- **Python 3.10 锁定**，numpy 锁 1.23.5，与新版 torch 常冲突。
- **无 flash-attn / triton 依赖** → 无相关编译坑。
- **ONNX 是流式实时推理的前置条件**（作者明说「需要将模型转为onnx」）。**帧率约束**：用 wenet 编码器须 **20fps**，用 hubert 须 **25fps**，否则特征对齐错误。
- **SyncNet 已被作者移除**：README 逐字「I removed SyncNet-related code and docs to keep the repo simpler.」→ 参照旧教程加 `--use_syncnet` 会失败。
- **推理默认跑 CPU**：[issue #74](https://github.com/anliyuan/Ultralight-Digital-Human/issues/74) 逐字：「推理时，使用文档中的语句：`python inference.py ...` 但是**全部用的是CPU去处理，GPU没动静**，是需要配置参数吗」→ 需自行改 device。
- 训练质量坑（[issue #18](https://github.com/anliyuan/Ultralight-Digital-Human/issues/18) 逐字）：「**200epochs是不是不太够**，训练了3个都有点抖」。
- 推理端音频特征提取慢（[issue #101](https://github.com/anliyuan/Ultralight-Digital-Human/issues/101) 逐字：「感觉推理部分做音频特征提取速度太慢了…有接入流式音频的可能性吗」）。

---

## 4. TMElyralab/MuseTalk

### (a) 统计口径（shields.io 原始值）
| 项 | 原始值 |
|---|---|
| stars | `{"label":"stars","message":"6.6k",...}` → **6.6k** |
| license | `{"label":"license","message":"not identifiable by github",...}` → **`not identifiable by github`**；但 `LICENSE` 文件首行逐字为 **「MIT License / Copyright (c) 2024 Tencent Music Entertainment Group」**（**实际是 MIT**） |
| last commit | `{"label":"last commit","message":"september 2025",...}` → 逐字 **`september 2025`**；Atom 精确值 **2025-09-26T05:44:17Z** → **已近一年无提交** |

### (b) 实现类型
**2D 视频驱动口型同步**：在 `ft-mse-vae` latent 空间做单步 inpainting，**不是 diffusion 采样**。官方 README 逐字：「modifies an unseen face according to the input audio, with a size of **face region of `256 x 256`**」；音频编码用冻结 `whisper-tiny`，图像用冻结 VAE，先验来自 `stable-diffusion-v1-4` 的 UNet 架构。

### (c) 显存需求 + 16GB 单卡能否跑 —— ⭐ 本项目有**同款硬件直接实测**
**① 🔴 [第三方实测·目标硬件直接命中] RTX 5060 Ti 16GB 上 ~3.6GB VRAM**
[MuseTalk issue #409](https://github.com/TMElyralab/MuseTalk/issues/409) 标题逐字：「**Field Report: MuseTalk V1.5 working on RTX 5060 Ti (Blackwell sm_120) with Python 3.12 + mediapipe patch**」（2026-03-18）。正文 Performance 段逐字：
> 「- **7sec audio → 30sec inference → MP4 output**
> - **~3.6GB VRAM for MuseTalk models**
> - Reference image face detection cached after first call」

文末环境逐字：「Setup: **RTX 5060 Ti 16GB | Ubuntu 24.04 | Python 3.12.3 | PyTorch 2.10.0+cu128 | MuseTalk V1.5**」。
同 issue 2026-08-05 用户 HeimdallCore 追加逐字：「We ran into the same Blackwell (RTX 5060 Ti, sm_120) situation... **We verified this end-to-end (real MuseTalk V1.5 inference, 361 frames, correct lip-sync in the output).**」

**② [官方自述] 最低 4GB VRAM（fp16）** README「Gradio Demo」逐字：「we tested the system on a Windows environment using an **NVIDIA GeForce RTX 3050 Ti Laptop GPU with 4GB VRAM**. In **fp16 mode, generating an 8-second video takes approximately 5 minutes**.」
**③ [第三方实测] 更老的卡也能跑（极慢）** [issue #392](https://github.com/TMElyralab/MuseTalk/issues/392) 逐字：「I was able to run musetalk with a **gtx 1060 3gb**. It's just **ultra slow**.」
**④ [第三方实测] 实时推理 11GB 足够，瓶颈是算力** [issue #310](https://github.com/TMElyralab/MuseTalk/issues/310) 用户 wanlichina 逐字：「**显存占用我测试下来，能做到11G**。也就是3080,4080,5080都能跑，但是**瓶颈不在显存，在GPU算力**，因为GPU占用率爆了，多开满足不了实时推理性能」；楼主未调 batch 时报 `Tried to allocate 400.00 MiB (GPU 0; 23.64 GiB total capacity; 22.15 GiB already allocated...)`，调到 `batch_size=15` 后「现在占用在20g左右」。
**⑤ [推算] fp16 权重体积** —— 对 HF 文件发 HTTP HEAD 读 `Content-Length`：`musetalkV15/unet.pth` = **3,400,074,924 B**（3.40GB ≈ 3.17GiB，即 **FP32**）→ **fp16 约 1.70GB**（[HF 仓库](https://huggingface.co/TMElyralab/MuseTalk)）。此为**推算值**。
**⑥ [官方自述] 训练显存 —— 16GB 单卡完全不可能训练**（基于 8× NVIDIA H20 测试）：

| Stage 1 | Batch Size | Grad Accum | Memory per GPU |
|---|---|---|---|
| | 8 | 1 | **~32GB** |
| | 32 | 1 | **~74GB** ✓ |

| Stage 2 | Batch Size | Grad Accum | Memory per GPU |
|---|---|---|---|
| | 1 | 8 | **~54GB** |
| | 2 | 8 | **~85GB** ✓ |

**→ 16GB 单卡结论**：
- **推理：完全够用，余量很大。** 目标硬件直接实测 **~3.6GB**；作者称 fp16 最低 4GB 可跑；第三方基准峰值 5.0–5.2GB。**16GB 单卡推理无压力。**
- **训练：单卡 16GB 不可行**（官方最低档 stage1 bs=8 需 ~32GB，stage2 需 ~54GB 起）。→ **16GB 只能用官方权重做推理，放弃自训练。**

### (d) 端到端延迟：首帧 + FPS
- `[官方自述]` **「real-time … 30fps+ on an NVIDIA Tesla V100」**（README 出现 3 次）。⚠️ 官方**只给了 V100**，无 3060/4090/sm_120 数据。
- `[推算·基于第三方实测]` **RTX 5060 Ti 16GB ≈ 5.8 fps（远未实时）**：由 issue #409 的「7sec audio → 30sec inference」推算，7s×25fps=175 帧 ÷ 30s ≈ **5.8fps**。⚠️ 该方法用批量推理而非 `realtime_inference.py` 流式脚本，未说明是否 fp16。
- `[第三方实测]` **OpenTalking 基准**（OmniRT 后端跑 MuseTalk，512×512/25fps）——

| Hardware | Steady FPS | First-turn total/ms | **TTFV/ms** | **Peak inference VRAM/GB** |
|---|---|---|---|---|
| RTX 3090 | **28.868** | 3235.518 | **1769.484** | **5.078** |
| RTX 4090 | **24.767** | 3605.564 | **2095.522** | **5.203** |

逐字引用：「Peak inference VRAM/GB | **5.078**」「TTFV/ms | **1769.484**」。来源：[OpenTalking · MuseTalk](https://datascale-ai.github.io/opentalking/latest/en/avatar_models/musetalk/)。**TTFV ≈ 1.77s（3090）/ 2.10s（4090），首轮总耗时 3.24–3.61s。** 该页自述逐字：「Steady FPS is model-generation throughput, not full user-perceived latency; STT, LLM, TTS, queueing, and WebRTC still affect the complete experience.」（**归属说明：OpenTalking 用 OmniRT 后端的结果，非原生 MuseTalk 脚本**。）
- 🔴 **[第三方实测] 官方 30fps+ 与社区 5–45fps 落差的技术根因**（[issue #33](https://github.com/TMElyralab/MuseTalk/issues/33)）：4090 用户逐字「**在不考虑这一步的情况下，实时性可以达到60+的fps。但是因为它的存在导致我们的性能只能在30fps左右**」——"这一步" = `musetalk/models/vae.py` 里 VAE decode 后的 `image.detach().cpu().permute(...)` 回传；后续解释逐字「不是拷贝慢，是因为 **gpu运算还未结束，会一直阻塞**」。→ **瓶颈在 VAE decode + GPU→CPU 同步，不在 UNet。**
- `[第三方]` **社区对「实时」的普遍质疑**：[#396](https://github.com/TMElyralab/MuseTalk/issues/396)（Has anyone actually been able to achieve real-time video generation?）、[#248](https://github.com/TMElyralab/MuseTalk/issues/248)（4090只能跑到5fps是怎么回事）、[#49](https://github.com/TMElyralab/MuseTalk/issues/49)、[#333](https://github.com/TMElyralab/MuseTalk/issues/333)、[#377](https://github.com/TMElyralab/MuseTalk/issues/377)、[#359](https://github.com/TMElyralab/MuseTalk/issues/359)。

### (e) 流式 / 可打断 —— 🔴 **有官方成员本人的否定确认（硬阻塞）**
- 🔴 [issue #33](https://github.com/TMElyralab/MuseTalk/issues/33)：用户 jinqinn 逐字「realtime-inference支持了图片的实时推理，但**不支持音频流失输入**，因musetalk引用了whisper组件…」；**项目组成员 czk32611 回复逐字：「不好意思，音频的流式处理我们暂时没有研究。」**
  → **官方成员确认：MuseTalk 无法边收音频边说，音频必须先完整给出。对"实时交互数字人"是硬阻塞**，也解释了为何必须靠 LiveTalking 改造包装。
- `[第三方]` **「实时」口径问题**：[issue #310](https://github.com/TMElyralab/MuseTalk/issues/310) 楼主逐字「实时推理的好像只有图片是实时生成的，**音频是最后生成出来的**」；维护者 zzzweakman 回复逐字「是的，因为代码里有**合成视频这一步**，会将声音和图像序列合成视频」→ **音视频合成是事后步骤，不是逐帧流式**。
- **可打断（barge-in）：未查到/无**；该能力由 LiveTalking 等上游框架封装提供。

### (f) 中文 / 半身-全身
- **中文：官方明确支持多语言。** README 逐字：「supports audio in various languages, such as **Chinese, English, and Japanese**」。
- **半身 / 全身**：模型**只改 256×256 面部区**，画面范围完全由你给的输入视频决定。README 逐字：「supports modification of the **center point of the face region** proposes, which **SIGNIFICANTLY** affects generation results」。**全身需靠上游框架**（如 LiveTalking 的「全身视频拼接」），README 未提全身。

### (g) 已知坑
- 🔴 **坑 1（最关键）：sm_120 必须 cu128，官方推荐的 cu118 直接报错。** [issue #409](https://github.com/TMElyralab/MuseTalk/issues/409) 逐字对照表：

  | PyTorch | CUDA | sm_120? |
  |---|---|---|
  | 2.6.0+cu124 | 12.4 | ❌ **no kernel image** errors |
  | 2.10.0+cu126 | 12.6 | ❌ Same errors |
  | **2.10.0+cu128** | **12.8** | ✅ **Works!** |

  正文逐字：「MuseTalk recommends PyTorch 2.0.1+cu118, but Blackwell GPUs need **cu128+ for native sm_120 kernel support**.」修复命令：`pip install torch==2.10.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128`。
- 🔴 **坑 2：mmpose/mmcv 在 Python 3.12 上装不上。** issue #409 逐字：「mmcv has **no pre-built wheels for Python 3.12** on any CUDA index. **Building from source fails due to `pkg_resources` removal in Python 3.12**.」
- 🔴 **坑 3：mmcv 新版仍有 C++ 编译不兼容。** 2026-08-05 HeimdallCore 逐字：「mmcv (we tried both **2.0.1** and the latest **2.2.0**) fails to compile against current PyTorch with an **ambiguous-overload C++ error around `Float8_e5m2fnuz`** operators — a genuine upstream incompatibility, not fixable with a flag.」其干净方案逐字：改用独立 **`face-alignment`** PyPI 包（MIT，FAN-based），「no MuseTalk dependency, no CUDA/C++ build step, and it returns native **68-point dlib-format** landmarks directly」，仅需改 2 处（`musetalk/utils/preprocessing.py`），**并已端到端验证 361 帧**。→ **对 5060 Ti + Py3.12 的部署建议：直接采用 face-alignment 方案，绕开 mmcv 编译。**
- 🔴 **坑 4：50 系实操记录**（[issue #334](https://github.com/TMElyralab/MuseTalk/issues/334)）：楼主逐字「现在不支持50系列显卡，希望可以兼容一下」；wanlichina 逐字「可以兼容的，自动动手装最新的 **pytorch 2.7+cu128**，修改模型的 `torch.load` 的加载参数」（**新版 PyTorch 默认 `weights_only=True` 会拒绝加载旧 checkpoint**）；后续 `mim install mmcv==2.0.1` 在 cu128 编译失败，修正逐字「**MMCV 2.0.1不支持cu128，你改成 mim install mmvc==2.1.0 就好了**」；codestart-zhu 逐字确认「环境是 **pytorch2.7+cu128，现在5090能跑起来的**」。
- **坑 5：推理未加 no_grad 导致显存暴涨**（与 LiveTalking 同源）：[MuseTalk PR #349](https://github.com/TMElyralab/MuseTalk/pull/349) 标题逐字「Fix issue #235: Use `torch.no_grad()` in inference to prevent excessive…」（已 merged）。
- **坑 6：依赖陈旧且有 tensorflow。** `requirements.txt` 逐字：`diffusers==0.30.2`、`accelerate==0.28.0`、`numpy==1.23.5`、**`tensorflow==2.12.0`**、`transformers==4.39.2`、`gradio==5.24.0`。运行时会打印 `TF-TRT Warning: Could not find TensorRT`（**噪音非错误**）。**无 flash-attn、无 triton、无 onnx/tensorrt 依赖** → 不存在这些编译坑。
- ⚠️ **上游停滞**：仓库最后提交 2025-09-26；**其 2026 年的全部 Blackwell 适配案例（#409/#334）都是社区自发、尚未合并进主干** → 选它就得自己维护补丁分支。

---

## 5. bytedance/LatentSync

### (a) 统计口径（shields.io 原始值）
| 项 | 原始值 |
|---|---|
| stars | `{"label":"stars","message":"6.1k",...}` → **6.1k**；`[第三方]` 精确 **6,090** |
| license | `{"label":"license","message":"Apache-2.0","color":"green",...}` → **Apache-2.0**（`LICENSE` 实取到 Apache License 正文，确认） |
| last commit | `{"label":"last commit","message":"june 2025",...}` → 逐字 **`june 2025`**；Atom 精确值 **2025-06-20T07:36:51Z** → **停更 15 个月** |

### (b) 实现类型
**2D 视频驱动、音频驱动的 diffusion 口型同步**（latent-space，非 3D、非 3DGS）。版本沿革：**1.5（2025-03-14）**、**1.6（2025-06-11，改为 512×512 训练以缓解模糊）**。GitHub Releases 为空，最新版即 1.6。

### (c) 显存需求 + 16GB 单卡能否跑 —— **官方 1.6 门槛超 16GB，但第三方在 16GB 上实测跑通（压线）**
**① [官方自述] 推理最低显存**（README 逐字）：
> 「Minimum VRAM for inference:
> - **8 GB** with LatentSync 1.5
> - **18 GB** with LatentSync 1.6」

→ **1.6 把门槛从 8GB 抬到 18GB，对 16GB 卡是「退步」。**
**② [官方自述] 训练显存**（逐字）：「`stage1.yaml`: … requires **23 GB**」「`stage2.yaml`: … **30 GB**」「`stage2_efficient.yaml`: … **20 GB**」「`stage1_512.yaml`: … **30 GB**」「`stage2_512.yaml`: … **55 GB**」。
**③ 🔴 [第三方实测·关键] RTX 5070 Ti 16GB 上跑通，eager 模式峰值 15,807MB（贴天花板 96%）**
[issue #365](https://github.com/bytedance/LatentSync/issues/365)（opened Jul 3 2026，标注 "directed, tested, and validated by Petrus Vermaak"），测试系统 **RTX 5070 Ti 16 GB**，夹具 9.68s 输出 / 242 帧 / 512 face / 20 步 / guidance 1.5：

| 配置 | 耗时 | 每输出秒耗时 | 相对基线 | **Peak total VRAM** |
|---|---|---|---|---|
| `exact20` PyTorch eager | 397.512 s | 41.065 s | baseline | **15,807 MB** |
| `optimized20` TensorRT-RTX exact | 268.160 s | 27.702 s | **32.54% faster** | **12,551 MB** |

→ **16GB 能压进去（裸跑贴天花板 96%，几乎无余量），TensorRT 后留余量。**
**④ [第三方] 旁证 16GB 吃紧**：[issue #335](https://github.com/bytedance/LatentSync/issues/335) 中文标题逐字「**因為顯存不夠，只有16GB，想詢問是否能使用v1.5的版本？**」；[#314](https://github.com/bytedance/LatentSync/issues/314)「8G显存运行不了1.5」；[#284](https://github.com/bytedance/LatentSync/issues/284)「v1.6 是否可以在8G vRam 运行？」；[#278](https://github.com/bytedance/LatentSync/issues/278)「4090 24G train stage1 OOM」。
**⑤ [第三方] 权重体积**（[ComfyUI-LatentSyncWrapper README](https://raw.githubusercontent.com/ShmuelRonen/ComfyUI-LatentSyncWrapper/main/README.md) 逐字）：`latentsync_unet.pt (~5GB)`、`stable_syncnet.pt (~1.6GB)`；该封装自称「Optimized to run on **20GB VRAM (RTX 3090 compatible)**」（**与"Optimized"表述自相矛盾，仅量级参考**）。

**→ 16GB 单卡结论**：⚠️ **官方口径不达标（1.6 需 18GB），但实测可压进去（15,807MB 裸跑 / 12,551MB TensorRT）**；**1.5 版本官方门槛 8GB，稳妥可用。训练：16GB 完全不可行（20–55GB）。**
🔴 **但真正的问题不是显存而是速度**（见 (d)）。

### (d) 端到端延迟：首帧 + FPS
- 🔴 `[第三方实测]` 由上表换算：**≈0.61–0.90 帧/秒**，即 **27.7–41.1 秒墙钟换 1 秒成片 —— 比实时慢 28~41 倍**。
- **首帧延迟：未查到（且不适用——离线批处理）。**
- `[第三方]` 旁证 [issue #137](https://github.com/bytedance/LatentSync/issues/137)：4.17 it/s、**GPU 利用率仅 1%-3%**、「GPU Memory Usage: 6GB/ 16GB」（1.5 时代）→ **显存不是瓶颈，算力/利用率才是**。

### (e) 流式 / 可打断
**❌ 均不支持。** 官方只提供 `python gradio_app.py` 与 CLI，属**离线批处理**视频生成。`[第三方]` [issue #326](https://github.com/bytedance/LatentSync/issues/326) 标题即「How to run inference on longer video?」（长视频需自行切段）。无 barge-in/双工。

### (f) 中文 / 半身-全身
- **中文：官方专门优化过。** README 1.5 发布说明逐字：「improves performance on **Chinese videos**」。
- **半身 / 全身：🔴 不支持。** 管线是 **InsightFace landmark → 仿射对齐裁脸 → 256×256 / 512×512**，**只覆盖面部区域**；README 无半身/全身声明。

### (g) 已知坑
- 🔴 **坑 1：torch 钉在 2.5.1 + cu121 —— 不含 sm_120 内核。** `requirements.txt` 逐字：`torch==2.5.1`、`torchvision==0.20.1`、`--extra-index-url https://download.pytorch.org/whl/cu121`。按 [PyTorch 2.7 发布说明](https://pytorch.org/blog/pytorch-2.7/)，Blackwell 支持从 **2.7 + cu128** 才引入 → **照装会在 5060 Ti 上撞 `no kernel image is available`，必须自行升到 torch ≥2.7 + cu128/cu13x**，而任何 torch 升级都可能打破下面这批老旧 pin。
- ✅ **好消息：requirements 里没有 xformers、也没有 flash-attn** → **无 flash-attn/triton 现场编译 sm_120 的坑**（比多数同类项目友好）。
- **其余 pin**（逐字）：`onnxruntime-gpu==1.21.0`、`insightface==0.7.3`、`face-alignment==1.4.1`、`numpy==1.26.4`、`gradio==5.24.0`、`DeepCache==0.1.1`、`diffusers==0.32.2`、`transformers==4.48.0`、`decord==0.6.0`、`mediapipe==0.10.11`。
  - 🔴 **`onnxruntime-gpu==1.21.0` 与目标机 CUDA 13.2 跨大版本**：ORT 官方文档逐字「**1.26.x-1.21.x | CUDA 12.8 | cuDNN 9.x | Default GPU package build before 1.27**」，且逐字「ONNX Runtime built with CUDA 12.8 are compatible with any **CUDA 12.x** version; ONNX Runtime built with CUDA 13.0 are compatible with any **CUDA 13.x** version」→ **1.21.0 是 cu12.8 构建，只保证与 CUDA 12.x 兼容，而宿主是 CUDA 13.2**。`[推断/风险]` **要么补 CUDA 12 运行库（如 `nvidia-cuda-runtime-cu12`，否则可能缺 `libcudart.so.12`），要么升 ORT ≥1.27（cu13.0）后重测。**（Blackwell 专项：搜索索引显示 ORT 有 PR #23928「Extend CMAKE_CUDA_FLAGS with all Blackwell compute capacity」，说明 ORT 曾专门补 Blackwell arch 编译标志，**但该 PR 进入哪个版本未能核实**。）
  - ✅ **换 torch 的连带脆点之一是好消息**：ORT 文档逐字「PyTorch 2.3 uses cuDNN 8.x, while **PyTorch 2.4 or later uses cuDNN 9.x**」→ **2.5.1 与 2.7 同为 cuDNN 9，不冲突**。其余脆点：`decord==0.6.0` 新 Python 常无轮子、`insightface==0.7.3` 需现场 C++ 编译、`DeepCache==0.1.1` + `diffusers==0.32.2` 组合作者从未验证（2025-06 后停更）。
  - `insightface==0.7.3`：需 C++ 编译（无预编译 wheel 的常见坑）；**未查到**本项目针对 sm_120 的 insightface 编译报告。
- `[第三方]` [issue #270](https://github.com/bytedance/LatentSync/issues/270)「No available kernel. Aborting execution.」（**未逐条核实是否 sm_120**）。
- 🔴 **上游基本无响应能力**：Issues 页显示 **「New issue — Issue creation is restricted in this repository」**；最后提交 2025-06-20，**15 个月无更新**，2026 年的 Blackwell 生态变化不会有官方适配。
- **未查到**是否存在 1.7 版本。

---

## 6. Rudrabha/Wav2Lip

### (a) 统计口径（shields.io 原始值）
| 项 | 原始值 |
|---|---|
| stars | `{"label":"stars","message":"13k",...}` → **13k**；`[第三方]` 精确 **13,219** |
| license | `{"label":"license","message":"not specified","color":"lightgrey",...}` → **`not specified`**。仓库**根目录无 LICENSE 文件**（`master`/`main` 两分支均取不到）。README 末尾逐字给出实际条款（见 (g)） |
| last commit | `{"label":"last commit","message":"june 2025",...}` → 逐字 **`june 2025`**；Atom 精确值 **2025-06-22T02:41:20Z** |

### (b) 实现类型
**2D 视频驱动口型同步**（CNN 编码器-解码器 + 空间注意力），audio-driven lip-sync 的鼻祖级实现（2020）。

### (c) 显存需求 + 16GB 单卡能否跑
**① [官方] 本仓库 README 完全没给显存要求**（连推荐 GPU 都没有），只给 `Python 3.6` + ffmpeg。**具体 fp16 权重体积 / 实测峰值显存：未查到**（权重走 Google Drive，仓库不发布文件）。
**② [他人估算]** LiveTalking 作者估算 **Wav2Lip ~200MB（FP32）**、输入 256×256（[来源](https://www.cnblogs.com/livetalking/articles/22134438)）。
**③ [作者实测]** LiveTalking 作者在 **4090** 上实测「wav2lip模型…**占用显存在1.3G，推理速度能达到750fps**」（[来源](https://blog.csdn.net/lipku/article/details/148593974)）。
**④ [作者自述]** LiveTalking 官方表：wav2lip256 在 **RTX 3060 上 60fps**、3080Ti 上 120fps，**「wav2lip256 推荐 RTX 3060 及以上」**。

**→ 16GB 单卡结论**：`[推断]` **显存侧毫无压力**（96×96 小卷积网，封装方实测 1.3GB），**16GB 跑得动几无悬念**。**但本仓库无官方/实测显存数字，瓶颈是速度、画质与软件栈（见 (g)），不是显存。**

### (d) 端到端延迟：首帧 + FPS
- **本仓库：时延数字未查到。** 唯一线索 [issue #584](https://github.com/Rudrabha/Wav2Lip/issues/584) 原文「both are **slow and very GPU consuming**... I am about to drop this technology and give up on it」——**只有投诉、无数字**。
- 经 LiveTalking 封装后的数字见上：**RTX 3060 = 60fps**（官方表）、**RTX 4090 = 750fps 推理速度 @1.3GB**（作者实测）。**首帧延迟：未查到。**

### (e) 流式 / 可打断
**本仓库：❌ 无（整段批处理）**：`python inference.py --checkpoint_path ... --face ... --audio ...` → 输出 `results/result_voice.mp4`。
**流式能力来自上游封装**：LiveTalking 明确支持「数字人说话被打断」+ WebRTC/RTMP 推流，并把 wav2lip 作为首选实时后端。

### (f) 中文 / 半身-全身
- **中文**：纯音频驱动（mel 频谱），README 逐字「Works for any identity, voice, and language」→ **语言无关**；但**训练集是 LRS2（英文）**，`[推断]` 中文质量弱于英文。
- **半身 / 全身**：**只改嘴部区域，身体姿态完全不动**（能"保留"构图但**无身体动作生成**）。**半身/全身：未查到专门声明**；全身由上游框架（LiveTalking 全身拼接）实现。

### (g) 已知坑 —— 🔴 **本报告风险最高的项目**
- 🔴🔴 **坑 1：`torch==1.1.0`（2020 年版本）+ `Python 3.6` 硬门槛。** `requirements.txt` 逐字：`torch==1.1.0`、`torchvision==0.3.0`、`librosa==0.7.0`、`numpy==1.17.1`、`numba==0.48`、`opencv-python==4.1.0.25`。**torch 1.1.0 根本不支持 Python 3.12，更不含 sm_120 内核**；Ubuntu 24.04 + CUDA 13.2 下无法照抄。
  **PyPI 硬证据**（[pypi.org/pypi/torch/1.1.0/json](https://pypi.org/pypi/torch/1.1.0/json)）：该版本全仓**仅 9 个 wheel**、上传日 **2019-04-30**、Python 标签**只有 cp27/cp35/cp36/cp37——没有任何 cp38+ 轮子**，文件名 `torch-1.1.0-cp37-cp37m-manylinux1_x86_64.whl`（**manylinux1 = 2010 年 ABI**）。同批：torchvision 0.3.0→2019-05-22、numpy 1.17.1→2019-08-27、numba 0.48→2020-06-30。**三层含义：① 装不上（Ubuntu 24.04 是 Py3.12，无 wheel，pip 必失败）② 跑不了（CUDA 9/10 时代 arch，必报 no kernel image）③ 生态全断（numba/numpy/librosa 在 Py3.12 均无轮子）。** → **该 requirements.txt 在 2026 年已是「不可复现文档」，只能当算法参考、用现代 torch 重写**（社区各 fork/ComfyUI 节点本质都这么做）。要在 5060 Ti 上跑必须**彻底替换依赖栈**（这已不是「装个环境」，而是「移植项目」）。
- ✅ **依赖无版本 pin**（无 torch/CUDA 锁定，README 只说 `pip install -r requirements.txt`）→ sm_120 兼容性**取决于你自装 torch**（`[推断]` 需 ≥2.7 + cu128，与 Duix #624 里「PyTorch 2.7+ added Blackwell support」说法一致）。
- 🔴 **坑 2：许可证禁止商用。** README 逐字：
  > 「All results from this open-source code or our demo website should only be used for **research/academic/personal purposes only**. As the models are trained on the LRS2 dataset, **any form of commercial use is strictly prohibited**. For commercial requests please contact us directly!」
  > 「## License and Citation
  > This repository can only be used for **personal/research/non-commercial purposes**.」
  → **完全不能用于商业数字人产品**，且仓库**无 LICENSE 文件**（默认全版权保留）。
- 🔴 **坑 3：事实性停更。** README **开头第一段已不是论文说明，而是 "Commercial Version … the commercial version is of a much higher quality than the old open source model!"** + `pip install syncsdk` 导流 sync.so，**论文被挤到尾部** → 2025-06-22 那次 push 的性质是**商业导流而非代码**。Issue 追踪器最新活动是 **#613 opened Jan 3, 2024**、**#584 opened Nov 10, 2023（至今 Open 无人答）**。
- **社区推荐的后继者**：可核实的血缘是 **LatentSync / MuseTalk**（LatentSync README 致谢逐字「Some code are borrowed from MuseTalk, StyleSync, SyncNet, **Wav2Lip**」）；作者本人的商业后继是 **sync.so**（`rudrabha@synclabs.so`）。**"社区公认排序"未查到权威依据，本报告不下结论。**
- **S3FD 人脸检测权重需手动下载**（README 自己给了备用 SharePoint 链接，说明主链常挂）。

---

## 7. KwaiVGI/LivePortrait → **已改名为 KlingAIResearch/LivePortrait**

> 🔴 **2026 状态变更**：`KwaiVGI/LivePortrait` 重定向为 **`KlingAIResearch/LivePortrait`**（仓库页 `<meta description>` 逐字：「Contribute to **KlingAIResearch/LivePortrait** development」）。
> ⚠️ **抓取注意**：该仓库说明文件是**小写 `readme.md`**，`README.md` 返回 **404**。

### (a) 统计口径
| 项 | 值 |
|---|---|
| stars | shields（KwaiVGI 路径）`{"label":"stars","message":"19k",...}` → **19k**；2026-09-22 重定向后页面 HTML `"stargazerCount":19084` → **19,084** |
| license | shields `{"label":"license","message":"not identifiable by github",...}` → **`not identifiable by github`**；`LICENSE` 文件首行逐字 **「MIT License / Copyright (c) 2024 Kuaishou Visual Generation and Interaction Center」** |
| last commit | shields `{"label":"last commit","message":"june",...}` → 逐字 **`june`**；Atom 精确值 **2026-06-01T17:24:54Z** → 约 3.7 个月无提交（稳定/低维护） |

### (b) 实现类型
**视频驱动（video-driven）肖像动画 —— 隐式 3D 关键点 + warping + SPADE 生成器。** ⚠️ **不是 3DGS**，也**不是音频驱动的 lip-sync**。
⚠️ **最容易误选的一点**：README 逐字定位：「efficient **portrait-animation** (humans, cats and dogs) solution adopted by major video platforms—Kuaishou, Douyin, Jianying, WeChat Channels」。**不能**单独做「输入音频→说话数字人」，必须配驱动视频/模板信号。

### (c) 显存需求 + 16GB 单卡能否跑
- **官方未给出推理显存数字**（readme 全文**无 VRAM 段落**）→ 官方口径：**未查到**。
- **官方权重规模**（`assets/docs/speed.md` 原文表格，表头口径为 **RTX 4090 + 原生 PyTorch + `torch.compile`**、单帧）：

| Model | Parameters(M) | Model Size(MB) | Inference(ms) |
|---|---|---|---|
| Appearance Feature Extractor | 0.84 | 3.3 | 0.82 |
| Motion Extractor | 28.12 | 108 | 0.84 |
| Spade Generator | 55.37 | 212 | 7.59 |
| Warping Module | 45.53 | 174 | 5.21 |
| Stitching and Retargeting Modules | 0.23 | 2.3 | 0.31 |

  原文小注逐字：「The values for the Stitching and Retargeting Modules represent the combined parameter counts and total inference time of three sequential MLP networks.」
  合计 **≈130.1M 参数 / ≈499.6 MB（FP32）** → `[推算]` **fp16 权重 ≈ 0.25–0.5GB**。
- `[第三方声明]` [Lytanshade/LivePortrait-pinokio](https://github.com/Lytanshade/LivePortrait-pinokio) 标题逐字：「Bring a portrait to life using a Video source -> **6GB VRAM ~8GB install**」（第三方封装声明，未见 nvidia-smi 截图）。
- `[第三方·官方渠道]` **aigcpanel 官方模型市场把 LivePortrait 包标为 **8G+ / 16G+** 且标「支持50显卡」**（[来源](https://aigcpanel.com/zh/asset?tag=视频换口型)）→ **这是目前唯一来自"分发方官方"的门槛数字**。

**→ 16GB 单卡结论**：**✅ 显存余量充足。官方无实测数字**；三条不同口径的独立证据为——官方权重 **499.6MB**、第三方 pinokio **6GB VRAM**、aigcpanel 官方包标 **8G+** → **建议取 8GB 作保守门槛（最保守的一条）**，16GB 均有充足余量。

### (d) 端到端延迟：首帧 + FPS
- `[官方自述]` **单帧推理 RTX 4090 + `torch.compile`**：`0.82+0.84+7.59+5.21+0.31 = ` **14.77 ms/帧** → `[推算]` **≈ 67.7 FPS**（[speed.md](https://github.com/KlingAIResearch/LivePortrait/blob/main/assets/docs/speed.md)）。
- `[官方自述]` **`torch.compile` 首次开销**逐字：「The first-time inference triggers an optimization process (**about one minute**), making subsequent inferences **20-30% faster**.」（→ **首帧延迟含约 1 分钟编译**。）
- **业务意义的首帧延迟：本项目语境不存在** —— 视频/模板驱动，**没有 ASR→TTS→口型链路**，故未定义、**未查到**。
- 相对锚点（官方逐字）：macOS Apple Silicon 走 MPS「this maybe **20x slower than RTX 4090**」。

### (e) 流式 / 可打断
- **❌ 官方无流式/双工声明。** `app.py` 是 Gradio 交互推理，`inference.py` 是离线视频生成 → **逐段/离线推理**。
- **实时能力来自第三方改造**（官方 README 收录）：[FasterLivePortrait](https://github.com/warmshao/FasterLivePortrait)（「Faster **real-time** version using **TensorRT**」）、**FacePoke**（逐字「A **real-time** head transformation app, controlled by your mouse!」）、**ComfyUI-AdvancedLivePortrait**（逐字「A faster ComfyUI node with **real-time preview**」）。
- `[第三方实测]` **FasterLivePortrait 的真实实时数字**（[warmshao/FasterLivePortrait](https://github.com/warmshao/FasterLivePortrait)，默认分支 `master`）逐字：「Achieved **real-time running of LivePortrait on RTX 3090 GPU** using TensorRT, reaching speeds of **30+ FPS** … **including pre- and post-processing**, not just the model inference speed.」→ **注意：官方的 67.7 FPS 是纯模型单帧，30+ FPS 才是含前后处理的端到端。**
- **可打断/barge-in：无**（无 ASR/LLM/TTS 环节）。

### (f) 中文 / 半身-全身
- **中文：与语言无关**（不涉及 ASR/TTS），驱动信号是视频/模板 `.pkl`，**无中文语义层**。
- **半身/全身**：官方定位 **portrait（人像/头肩）**，支持 **humans / cats / dogs** 三模式。**无半身/全身声明，全身未查到。**

### (g) 已知坑 —— 🔴 **onnxruntime 生态风险最集中**
1. 🔴 **`onnxruntime-gpu==1.18.0` 硬钉**。`requirements.txt` 逐字：`-r requirements_base.txt` / `onnxruntime-gpu==1.18.0` / `transformers==4.38.0`。**1.18.0 是 2024 年构建，不含 sm_120 内核** → Blackwell 上 `CUDAExecutionProvider` 不可用、**回退 CPU**（见第 0 节 Natfii 说明）。
2. 🔴 **sm_120 + RTX 5060 的 ORT 已知缺陷**：[microsoft/onnxruntime issue #27621](https://github.com/microsoft/onnxruntime/issues/27621) 标题逐字：「[CUDA EP] **Silent deadlock** when calling `InferenceSession.run()` from Python thread on **Blackwell GPU (RTX 5060, sm_120, Windows)**」。（Windows 复现，Linux 未确认，但同为 sm_120 + RTX 5060 平台，属**高风险预警**。）
3. **要用 Blackwell 必须离开官方 pin 的 1.18.0**：或升到官方 **≥1.27（CUDA 13.0 构建）**，或用第三方重建 wheel。
4. **torch 安装示例只到 CUDA 12.1**（readme 给 cu111 / cu118 / **cu121**），且原文警告逐字：「On Windows systems, some higher versions of CUDA (such as 12.4, 12.6, etc.) may lead to unknown issues. You may consider **downgrading CUDA to version 11.8** for stability.」→ **对 CUDA 12.8/13.x + Blackwell 官方零说明。**
5. **X-Pose（Animals 模式）需自编译 CUDA 算子**：readme 逐字：「You need to build an OP named **`MultiScaleDeformableAttention`** first … which is used by X-Pose」→ sm_120 下需自行 `TORCH_CUDA_ARCH_LIST=12.0` 编译。**Humans 模式可跳过**（readme 逐字：「The step of Check your CUDA versions is **optional** if you only want to run **Humans mode**」）。
6. **无 flash-attn / triton 依赖** → 少一类坑。
6b. 🔴 **若要上实时档（FasterLivePortrait），依赖地狱对 sm_120 更致命**（[warmshao/FasterLivePortrait](https://github.com/warmshao/FasterLivePortrait) 逐字）：
   - 「Install **TensorRT 8.x (versions >=10.x are not compatible)**」→ **必须 TensorRT 8.x，10.x 以上不兼容**；
   - 需**自编译 grid_sample TRT 插件**（5D 输入）；
   - 「The latest **onnxruntime-gpu still doesn't support grid_sample cuda**」→ 须从 microsoft/onnxruntime 的 `liqun/ImageDecoder-cuda` 分支**源码编译 onnxruntime_gpu 1.17.0**。
   → 在 CUDA 13.2 + sm_120 上同时满足「TRT 8.x」与「自编译 ORT 含 sm_120」**难度极高**，**大概率需要换老 CUDA 环境或放弃实时档**。
7. 🔴 **许可证的隐藏限制（商用重点）**：根 `LICENSE` = MIT，但**末尾附注逐字**：
   > 「The code of InsightFace is released under the MIT License.
   > **The models of InsightFace are for non-commercial research purposes only.**
   > If you want to use the LivePortrait project for commercial purposes, you should **remove and replace InsightFace's detection models** to fully comply with the MIT license.」
   → **代码本身 MIT 可商用，但默认人脸检测依赖 InsightFace 模型（仅限非商业研究）**；商用须替换检测模型（README 给出线索：`ComfyUI-LivePortraitKJ` 用 **MediaPipe 替代 Insightface**）。

---

## 8. HumanAIGC-Engineering/OpenAvatarChat

### (a) 统计口径
| 项 | 值 |
|---|---|
| stars | shields `{"label":"stars","message":"3.8k",...}` → **3.8k**；2026-09-22 页面 HTML `"stargazerCount":3773` → **3,773** |
| license | shields `{"label":"license","message":"Apache-2.0","color":"green",...}` → **Apache-2.0** |
| last commit | shields `{"label":"last commit","message":"july",...}` → 逐字 **`july`**；Atom 精确值 **2026-07-31T00:28:29Z** → 活跃 |
| 最新版本 | **0.6.0（2026.04）**；上一版 0.5.1（2025.08.19） |

### (b) 实现类型
**全链路实时数字人对话框架（ASR + LLM + TTS + Avatar）+ 多种神经渲染后端**。0.6.0 起支持 **LiteAvatar、LAM、MuseTalk、SoulX-FlashHead** 四种数字人技术。README 逐字：「模块化的交互数字人对话实现」「支持 **LiteAvatar、LAM、MuseTalk、FlashHead** 等多种数字人技术」。
- **LiteAvatar** = 2D 面部动画（2.5D）；**LAM** = 3D 表达参数（`aigc3d/LAM_Audio2Expression`）；**MuseTalk** = 视频驱动口型；**SoulX-FlashHead** = 扩散模型实时流式说话头（0.6.0 新增，逐字「基于扩散模型的**实时流式**说话头生成」）。

### (c) 显存需求 + 16GB 单卡能否跑
**① [第三方整理/官方旧版 README] 配置档显存表**（DeepWiki 标注数据源为 `README.md 254-345`；⚠️ **当前 main 的 README 已不含此表**，现为 7581 字节）：

| Configuration File | VRAM Required | Best For |
|---|---|---|
| `chat_with_minicpm.yaml` | **20GB+（int4 后 10GB）** | 全本地、全隐私 |
| `chat_with_lam.yaml` | **2-3GB** | 多会话、3D 数字人 |
| `chat_with_qwen_omni.yaml` | **2-3GB** | 语音到语音端到端 |
| `chat_with_openai_compatible.yaml` | **8-10GB** | 本地 TTS + 云端 LLM |

来源：[deepwiki: Getting Started](https://deepwiki.com/HumanAIGC-Engineering/OpenAvatarChat/2-getting-started)。

**② [第三方整理] 逐 Handler 显存分解**（[deepwiki: System Requirements](https://deepwiki.com/HumanAIGC-Engineering/OpenAvatarChat/1.1-system-requirements)）：

| Handler | VRAM Usage |
|---|---|
| ASR（SenseVoice） | **~500MB** |
| LLM（MiniCPM-o FP16） | **~20GB** |
| LLM（MiniCPM-o int4） | **~10GB** |
| TTS（CosyVoice） | **~2GB** |
| **Avatar（LiteAvatar）** | **~3GB per session** |
| **Avatar（MuseTalk）** | **~4GB** |
| **Avatar（LAM）** | **~1GB**（客户端渲染卸载） |

多会话规则逐字：「**Maximum concurrent sessions = floor(available_VRAM / 3GB)** for LiteAvatar on GPU」，且「When `concurrent_limit` exceeds available VRAM divided by 3GB, **out-of-memory errors occur**」。
⚠️ **口径冲突**：另一页 [deepwiki: Avatar Systems](https://deepwiki.com/HumanAIGC-Engineering/OpenAvatarChat/6-avatar-systems) 给 MuseTalk 单会话 **~6-8GB**（与上表 4GB 不一致）→ **MuseTalk 后端的确切 GB 未查到定论**。

**③ [第三方实测·带实测图] RTX 3060 6G 笔记本实测：LAM 3.1GB / LiteAvatar 5.3GB**
[53AI 转载《低显存福音！OpenAvatarChat开源！实测仅需4G显存…》](https://www.53ai.com/news/OpenSourceLLM/2025061631429)（2025-06-16）逐字：
> 「在搭载 **RTX3060 6G 的笔记本电脑上运行测试** 运行 **LAM数字人，仅占用 3.1G 显存**！…运行 **LiteAvatar数字人，也仅需 5.3G 显存**！」

同文依赖逐字：「N卡**4G以上显存** • N卡支持 **CUDA >= 12.4**」；测试方案 = SenseVoice + LLM API + CosyVoice API。
⚠️ **口径修正（重要）**：**LiteAvatar 实测 5.3GB > LAM 3.1GB**，与上表 DeepWiki 的「LiteAvatar 3GB/session」不符 → **本报告采用「LAM ≈3GB、LiteAvatar ≈5–6GB」，不沿用 3GB 单会话口径。**

**④ 系统级要求**（同上 DeepWiki）：VRAM **min 10GB / rec 20GB+**；RAM 16GB min / **32GB+ recommended**。

**→ 16GB 单卡结论**：
- `chat_with_lam`（2–3GB）/ `lite-avatar`（3GB/session）/ `chat_with_qwen_omni`（2–3GB）→ **✅ 可跑且余量大**（LiteAvatar 在 16GB 上理论可开 floor(16/3)=**5 路并发**）。
- `chat_with_openai_compatible`（8–10GB，本地 TTS + 云端 LLM）→ **✅ 可跑**，适合目标机。
- `chat_with_minicpm`（**全本地 LLM，20GB+**）→ **❌ 16GB 不可，除非用 int4 档（10GB）**。
- MuseTalk 后端（4GB 或 6-8GB，口径不一）→ **✅ 可跑**。

### (d) 端到端延迟
- `[官方自述]` README 唯一明确数字逐字：「**低延迟优化**：通过 VAD 检测、语音缓冲、帧率控制等机制优化，**平均响应时间仅 2.2 秒**」；英文版逐字：「~**2.2s average response time**」。
- `[第三方整理]` DeepWiki 明确测得条件逐字：「System latency (**user speech end → avatar speech start**) measured at **~2.2 seconds on i9-13900KF + RTX 4090**」。
- 输出帧率：DeepWiki 逐字「RGB video frames at configured FPS (typically **25fps**)」。
- **首帧延迟、推理 FPS：官方未给（未查到）。**

### (e) 流式 / 可打断 —— ⭐ 本项目最强项
0.6.0 更新逐字：「**所有数字人均支持手动打断和双工打断模式**」→ **全后端支持手动打断 + 双工（barge-in）打断**。这在本批 11 个项目中是**最完整**的（LiveTalking 支持打断但未见双工声明；MuseTalk/LatentSync/Wav2Lip/LivePortrait 自身都没有）。

### (f) 中文 / 半身-全身
- **中文：原生强支持。** ASR 用 **SenseVoice**（阿里，中文强），TTS 用 **CosyVoice**（README 组件表：`FunAudioLLM/CosyVoice`），LLM 走百炼（阿里云）OpenAI 兼容 API。模型下载默认优先 **ModelScope**，国内网络友好。
- **半身/全身**：**LAM 为 3D 数字人**（走表达参数，理论可驱动全身模型）；LiteAvatar 为 **2D 面部动画**；MuseTalk/FlashHead 为**头肩/半身视频口型驱动**。**明确的全身（full-body）声明：未查到。**

### (g) 已知坑
1. ✅ **对 Blackwell 有利**：官方快速开始页逐字：「本项目的运行依赖 CUDA，请确保本机 NVIDIA 驱动程序支持的 **CUDA 版本 >= 12.8**」→ 12.8 起支持 sm_120，本机 CUDA 13.2 满足。
2. ✅ **官方文档明确处理了 RTX 50 系**（[deepwiki: System Requirements](https://deepwiki.com/HumanAIGC-Engineering/OpenAvatarChat/1.1-system-requirements) 逐字）：「CUDA | >=12.4 | NVIDIA driver must support this version; **CUDA 12.8 used for RTX 50-series**」「**RTX 50-series support requires CUDA 12.8**」「The Docker image **Dockerfile.cuda128** packages CUDA 12.8 dependencies」「General Linux with **NVIDIA driver 575.64.03+** for RTX 50-series」。目标机驱动 610.43.02 **满足**。
3. 🟡 **flash-attn 需源码编译**：快速开始页逐字「需要编译安装的包（如 **flash-attn**）会自动使用 `--no-build-isolation` 并限制并行编译线程数」→ sm_120 下 flash-attn 源码编译**耗时长且易失败**（官方已缓解但仍是坑）。
4. 🔴 **PyTorch 版本落后**：DeepWiki 记录 `pyproject.toml` 钉 **PyTorch 2.4.1 + CUDA 12.4** → **该组合不含 sm_120 内核**，Blackwell 上须自行升级到 cu128/cu130 轮子，会与 `install.py` 的版本解析产生冲突。
5. **Python 版本窗口窄**：**≥3.11.7, <3.12**。
6. **依赖 `uv` + git lfs 子模块**（文档要求 `git lfs install`、`git submodule update --init --recursive --depth 1`），模型体积大。
7. **多会话硬上限**：LiteAvatar 每会话 3GB，`floor(VRAM/3GB)`；Linux 单机多 session 是 0.5.1 起才支持。
8. ⚠️ **未查到 sm_120 专项 issue**（GitHub API 限流，未能做 issue 关键词检索）→ **建议后续补验**。

---

## 9. Kedreamix/Linly-Talker

### (a) 统计口径
| 项 | 值 |
|---|---|
| stars | shields `{"label":"stars","message":"3.5k",...}` → **3.5k**；页面 HTML **3,453** |
| license | shields `{"label":"license","message":"MIT","color":"green",...}` → **MIT**（`LICENSE` 首行「MIT License / Copyright (c) 2024 Kedreamix」确认） |
| last commit | shields `{"label":"last commit","message":"february",...}` → 逐字 **`february`**；Atom 精确值 **2026-02-10T05:17:02Z** → **已 7 个多月无提交** |

### (b) 实现类型
**多后端整合型对话数字人系统**（ASR + LLM + TTS + 可切换 Avatar 后端）：**Wav2Lip / Wav2Lipv2 / SadTalker / ER-NeRF / MuseTalk**。README 功能清单逐字：「Talker模型多选择：**Wav2Lip / Wav2Lipv2 / SadTalker / ERNeRF / MuseTalk** / Coming Soon」。属 2D 视频驱动为主 + ER-NeRF 神经渲染。

### (c) 显存需求 + 16GB 单卡能否跑
- **官方 README：训练与推理均未给出任何显存数字 → 官方口径「未查到」。**
- `[第三方估算·可信度存疑]` [CSDN《Linly-Talker对显卡配置的要求及性价比推荐》](https://blog.csdn.net/weixin_42561464/article/details/156119020) 逐字（以 FP16 为例）：
  > 「大语言模型（LLM） Llama-3-8B（INT4量化） **~6–8 GB** ｜ 若未量化，FP16 下需约 **16GB**」
  > 「语音合成（TTS） VITS / YourTTS **~2–4 GB**」「语音识别 Whisper-tiny / base **~1–2 GB**」
  > 「哪怕是最轻量化的组合（INT4量化LLM + TTS + Wav2Lip），总显存需求也轻松突破 **10GB**」
  > 「因此，**12GB 显存应被视为当前运行 Linly-Talker 的最低推荐门槛**」
  ⚠️ **该文风格接近 AI 生成长文、无 nvidia-smi 截图，数字为「模块估算求和」而非实测 → 只能标【第三方估算】，不可当实测引用。**
- 上游引用：README 转录 MuseTalk 官方声明「30帧每秒以上…**NVIDIA Tesla V100**」。

**→ 16GB 单卡结论**：**⚠️ 依赖你选的组合，无官方数字，需实机验证。** 若采用「int4 LLM + CosyVoice + Wav2Lip/MuseTalk」的轻档，按第三方估算约 10–12GB → `[推断]` **16GB 可跑**；若 LLM 用 FP16（约 16GB）则**不可**。README 亦逐字提示为省显存「改进的WebUI在**默认设置下不加载 LLM 模型**」。

### (d) 端到端延迟：首帧 + FPS
- **主仓：无端到端延迟数字。** 可引用：MuseTalk 在 **V100 上 >30 FPS**（转述上游）；README 逐字「加入了 MuseTalk 的方式…**速度基本能够达到实时的要求**」（**无具体 ms**）。
- **Linly-Talker-Stream（分仓）**：README 逐字「⚡ **Low-latency transport**: real-time audio/video transmission via **WebRTC**」→ **只有定性「低延迟」，无具体 ms（未查到）**。
- **首帧延迟：两仓库均未给出，未查到。**

### (e) 流式 / 可打断
- **主仓：非流式。** Gradio 轮次式；README 自述逐字「若考虑实时对话，可能需要**换个框架**，或者对 Gradio 进行魔改」。
- **分仓 [Linly-Talker-Stream](https://raw.githubusercontent.com/Kedreamix/Linly-Talker-Stream/main/README.md)：✅ 支持打断 + 全双工**，逐字「**Barge-in and interruption support**: more natural conversational rhythm」「full-duplex conversation experience」；架构逐字引用 LiveTalking 做「**streaming pipeline refactor**」。

### (f) 中文 / 半身-全身
- **中文：原生中文项目，支持强。** 中文 README 为一等公民；ASR 集成阿里 **FunASR**/OmniSenseVoice，TTS 用 **CosyVoice**（README 逐字「支持中文、英语、日语、粤语和韩语等多种语言的语音合成」）。
- **半身/全身：无全身支持。** SadTalker / Wav2Lip / ER-NeRF / MuseTalk 均为**头肩/上半身肖像口型驱动**；README 无全身声明 → **全身：未查到/不支持**。

### (g) 已知坑
- 🔴 **最大风险是维护停滞**：主仓最后提交 **2026-02-10**，**7 个多月无提交** → 2026 年的 CUDA/torch 生态演进（尤其 sm_120）**基本不会得到官方适配**。
- **sm_120/Blackwell 兼容性：官方零说明，未查到。** README 未声明 CUDA 版本要求；项目基于 2024–2025 技术栈（SadTalker 2023 / Wav2Lip 2020 / CosyVoice 早期档），实际部署需自行把 torch 升到 cu128/cu130。
- **SadTalker/Wav2Lip 仍在，但已非实时主力**：checkpoints 结构仍含 `SadTalker_V0.0.2_256.safetensors`、`wav2lip.pth`、`wav2lip_gan.pth`；但 2026 年 WebUI 里承担实时对话的是 **MuseTalk**。**此判定基于 README 描述，未逐行核验代码默认值。**
- **模型分散需多源下载**：checkpoints 需从 HuggingFace / ModelScope / 百度网盘分别下载（README 提示百度网盘默认命名 `sadtalker` 需重命名为 `checkpoints`），国内下载易中断。

---

## 10. xszyou/Fay

### (a) 统计口径
| 项 | 值 |
|---|---|
| stars | shields `{"label":"stars","message":"14k",...}` → **14k**；页面 HTML **13,545** |
| license | shields `{"label":"license","message":"GPL-3.0","color":"orange",...}` → **GPL-3.0**（`LICENSE` 首行逐字「GNU GENERAL PUBLIC LICENSE / Version 3, 29 June 2007」确认） |
| last commit | shields `{"label":"last commit","message":"yesterday",...}` → 逐字 **`yesterday`**；Atom 精确值 **2026-09-21T08:56:28Z** → **极活跃（核实前 1 天仍有提交）** |

### (b) 实现类型
**不是渲染引擎，而是「agent 框架 + 数字人连接器」。** README 逐字：「**Fay数字人框架，向上适配各种数字人模型技术，向下接入各式大语言模型**…为单片机、app、网站提供全面的数字人应用接口」。仓库 `<meta description>` 逐字：「fay是一个帮助数字人（**2.5d、3d、移动、pc、网页**）或大语言模型（openai兼容、deepseek）连通业务系统的 agent 框架」。README 有独立章节「使用数字人（**非必须**）」。
生态侧可核实的 UE 通道：[`xszyou/fay-ue5`](https://raw.githubusercontent.com/xszyou/fay-ue5/main/README.md) 独立仓库存在。**xuniren / metahuman 的官方名称级原文：本次未抓到，未查到**（README 只给飞书 wiki 跳转）。

### (c) 显存需求 + 16GB 单卡能否跑
- **Fay 本体：0 显存需求。** 官方 README 全文**无 GPU/显存要求**（环境段仅「Python 3.12」「Windows、macos、ubuntu」「ubuntu需要先安装gcc及portaudio」）。
- **GPU 取决于你外挂的模型**：ASR/TTS/LLM 走云端 API 时 **0 显存**；外挂本地 FunASR/CosyVoice + 本地 LLM 才占显存。
- `[第三方]` 官方镜像（Compshare）页面可核实：镜像名「数字人 **Fay数字人-3.1.1**」、**镜像大小 40 GB**、累计部署 255 次、最近更新 2025-12-05、**CUDA 版本栏为空** → [compshare 镜像页](https://www.compshare.cn/images/compshareImage-1cft3sk9gvta?ytag=GPU_fay)。

**→ 16GB 单卡结论**：**✅ 本体无 GPU 需求，必然"能跑"。** 但**端到端显存完全取决于你接的数字人模型与本地 ASR/TTS/LLM 组合**，Fay 官方未给任何数字，**未查到**。

### (d) 端到端延迟：首帧 + FPS
**官方无任何毫秒/FPS 数字（未查到）**；仅定性声明「**全时流式的支持**」「支持唤醒及**打断**对话」。

### (e) 流式 / 可打断
**✅ 支持。** README 逐字：「全时流式的支持」「支持唤醒及**打断**对话」「支持多用户多路并发」「支持服务器及单机模式」。这是其相对 LivePortrait / Linly-Talker 主仓的显著优势。

### (f) 中文 / 半身-全身
- **中文：原生中文项目，支持强。** ASR 集成 **FunASR**（README 致谢逐字「FunASR - 提供语音识别（ASR）能力」）。
- **半身/全身：由外挂数字人决定。** README 定位涵盖 **2.5D 与 3D**；3D 侧可通过 UE 连接器实现全身。**Fay 自身不含渲染，半身/全身 = 你所接数字人模型的能力**（具体全身支持**未在官方原文确认，未查到**）。

### (g) 已知坑
- 🔴 **GPL-3.0 许可证 = 商用传染性风险（重点）**：README 自称「完全开源，**商用免责**」，但 `LICENSE` **实为 GPL-3.0**。GPL-3.0 的 **copyleft 传染性**意味着：若你将 Fay 与其代码**结合分发**，衍生作品需以 GPL 兼容方式开源；「商用免责」是作者的口头承诺，**不是许可证授予的额外权利**。（**GPL-3.0 非 AGPL**，纯 SaaS/网络服务场景**不触发**网络条款，这点优于 AGPL；但**分发二进制仍受限**。）商业闭源集成前必须做法律评估。
- **sm_120/Blackwell：🟢 风险最低。** Fay 本体是纯 Python 编排层，**只要外挂组件不编译 CUDA 算子就不受 sm_120 影响**。
- **环境**：Ubuntu 需 `build-essential` + `portaudio19-dev`；Python 3.12。
- **配置门槛**：快速启动需从 `system.conf.bak` 生成 `system.conf` 并填 API key；默认公共资源**速度非常慢**（作者自述）。

---

## 11. modstart-lib/aigcpanel

### (a) 统计口径
| 项 | 值 |
|---|---|
| stars | shields `{"label":"stars","message":"5.6k",...}` → **5.6k**；页面 HTML **5,574** |
| license | shields `{"label":"license","message":"Apache-2.0","color":"green",...}` → **Apache-2.0**（`LICENSE` 首行确认） |
| last commit | shields `{"label":"last commit","message":"last friday",...}` → 逐字 **`last friday`**；Atom 精确值 **2026-09-18T04:46:10Z** → 活跃 |
| 技术栈 | README badge 逐字：`Framework-**TS+Vue3+Electron**` |

### (b) 实现类型
**不是自己的渲染引擎，而是 Electron 桌面客户端 + 本地模型运行时管理器（兼引擎集成商）。** README 逐字：
> 「AIGCPanel 是一款简单易用的一站式 AI 数字人**桌面应用**，支持 **Windows / macOS / Linux** 三平台。」
> 「软件内置**模型市场，支持一键下载启动包**，开箱即用；同时兼容**远程 API 模型**。」
> 「**数字人合成**：基于多种开源口型同步模型（**MuseTalk / LatentSync / Wav2Lip / Heygem**），将任意音频与人物视频精准对齐」
> 「**云端 AI 模型服务（无需本地显卡）（VIP）**」

→ **前端是 Electron 壳，真正的推理是它拉起的「本地模型启动包」或云端 API。** 对选型的意义：**一个把 MuseTalk/LatentSync/Wav2Lip/Heygem 统一封装的入口**。

### (c) 是否需要 GPU / 显存要求
官方文档给了最清晰的答案（[aigcpanel 官方文档](https://aigcpanel.com/zh/document/74)，逐字）：
> 「**界面软件本身对硬件无特殊要求，普通电脑/笔记本即可运行。AI 模型运行需要依赖 专业显卡 GPU。如果没有 NVIDIA 显卡，可以使用 AI 云模型 服务，无需本地显卡即可使用全部功能（云模型按使用量计费）。**」
> 「目前我们提供的模型大部分基于 **NVIDIA CUDA** 框架，如需本地运行，务必准备 NVIDIA 显卡。」

- **面板本体显存需求 = 0。**
- 🔴 **官方「模型市场」已逐包公布 GPU 门槛（本报告最硬的 aigcpanel 证据）** —— 来源：官方资源页 [aigcpanel.com/zh/asset?tag=视频换口型](https://aigcpanel.com/zh/asset?tag=视频换口型)（共 12 个资源，逐字）：

| 模型包 | 官方 GPU 门槛 | 50 系标注 | 备注 |
|---|---|---|---|
| Wav2Lip | **4G+ / 8G+** | — | |
| SadTalker | **4G+ / 8G+** | — | |
| **MuseTalk** | **6G+ / 16G+** | ✅ **支持50显卡** | 2026-09-16 v1.3.0；「V100 上 30fps+ 实时推理」 |
| Wav2Lip384 | 8G+ / 16G+ | — | VIP |
| **Heygem-v2** | **8G+ / 16G+** | ✅ **支持50显卡** | |
| LatentSync | **8G+ / 16G+** | — | |
| **LatentSync15** | **8G+ / 16G+** | ✅ **支持50显卡** | |
| **SoulX-FlashHead Lite** | **8G+ / 16G+** | ✅ **支持50显卡** | 「消费级 GPU 可达 **96 FPS** 实时推理」 |
| **LivePortrait** | **8G+ / 16G+** | ✅ **支持50显卡** | |
| LatentSync ComfyUI | 16G+ / 16G+ | — | |
| Heygem | 16G+ / 16G+ | — | VIP/SVIP |
| **FlashTalk（SoulX-FlashTalk-14B）** | **24G+ / 32G+** | — | **唯一超过 16GB 的包** |

→ **12 个包中 11 个官方门槛 ≤16GB；6 个包明确标注「支持50显卡」**（MuseTalk / LatentSync15 / SoulX-FlashHead Lite / LivePortrait / Heygem-v2 / FlashTalk）。另 LatentSync 文案称「大幅降低训练显存至 **20GB**」。

**→ 16GB 单卡结论**：**面板本身不吃显存，16GB 足够；且官方 12 个模型包中 11 个门槛 ≤16GB → 可按需选包。** 这同时印证了 **aigcpanel 确实会本地拉起引擎（不只是"调用"）**，GPU 需求完全由所选模型包决定。
⚠️ **但需注意各上游引擎的真实状态**：它集成的 **Wav2Lip 明确非商用**、**Duix/Heygem 在 sm_120 上仍有 [issue #624](https://github.com/duixcom/Duix-Avatar/issues/624) 的内核报错（尽管 aigcpanel 标了「支持50显卡」——两者口径冲突，建议实测）**。

### (d) 端到端延迟：首帧 + FPS
**未查到。** 定位是**离线视频合成工具**（「上传形象视频，输入文本或音频，一键生成换口型数字人视频」），README 与文档**均未给出延迟或 FPS**。「智能直播（VIP）」只声明弹幕实时监控（抖音/哔哩哔哩/虎牙/斗鱼/快手），**未声明画面渲染延迟**。

### (e) 流式 / 可打断
**未查到流式/打断能力。** 核心交互是**任务式（提交→生成→下载）**：README 逐字「任务列表可批量下载」「支持**断点续跑**、节点级状态追踪、运行历史查看」→ 更接近**批处理/半自动**，非双工实时对话。

### (f) 中文 / 半身-全身
- **中文：原生中文产品**（全中文 UI/文档），TTS/ASR 均含中文模型 → 中文**强**。
- **半身/全身**：口型同步基于**用户上传的人物视频** → **画面范围完全取决于素材**（上传全身视频即全身）。官方**无专门「半身/全身」功能声明，未查到**。支持**绿幕视频形象**与形象模板管理。
- 附加能力：**25+ 音视频工具箱**、内置模型市场（一键下载启动包）、Pro 版可视化工作流编排。

### (g) 已知坑
- **sm_120/Blackwell：官方未声明 CUDA/torch 版本要求（未查到）。** 由于它把引擎打包为「模型启动包」（内部可能含预编译 torch/ONNX），**Blackwell 上是否开箱可用取决于各启动包内的 torch/onnxruntime 是否含 sm_120 内核 → 属不可控的打包风险，建议逐个模型包实测**。
- **闭源打包**：模型启动包体积大、更新依赖官方分发节奏。
- **「云端 AI 模型服务」是 VIP 付费项**（按使用量计费）；免费路径需本地 NVIDIA 显卡。
- **Apache-2.0 是商用友好的**（无 GPL 传染问题）——**但注意：其集成的 MuseTalk / Wav2Lip / LatentSync / Heygem 各自的许可证与商用条款需单独核查**（例如 **Wav2Lip 明确非商用**、**Duix 有月活/用户数门槛**、**LivePortrait 须替换 InsightFace**）；aigcpanel 的 Apache-2.0 **不覆盖**第三方模型的权重许可。

---

## 12. 附加发现：本批次外的相关继任/配套项目（2026 状态）

| 项目 | stars | license | last commit | 定位与价值 |
|---|---|---|---|---|
| [anliyuan/FeatherTalk](https://github.com/anliyuan/FeatherTalk) | **93**（shields 逐字 `93`） | Apache-2.0 | **2026-09-04**（shields `september`） | **Ultralight-Digital-Human 的官方继任者**；视觉模型 5.46M 参数 / 144×144；FeatherHuBERT 音频算力降 40×；**C++/MNN，可纯 CPU 运行**；Apache-2.0。⚠️ 其 README 抓取返回 0 字节，显存/延迟数字**未查到** |
| [duixcom/Duix-Mobile](https://github.com/duixcom/Duix-Mobile) | **8.3k / 8,252** | `not identifiable by github` | **2026-08-05** | Duix 移动端实时 SDK（Android/iOS/Pad/车机/VR/IoT）；README 称 **<1.5s 延迟 / 手机端 120ms / 支持流式 barge-in**；**Duix 组织内 2026 最活跃** |
| [duixcom/duix-skills](https://github.com/duixcom/duix-skills) | 4 | `not identifiable by github` | 2026-08-12 | ⚠️ **云端 API skill**（要 `DUIX_API_KEY` + 预付费积分 + 订阅），**不是本地 16GB 部署后继者** |
| [Kedreamix/Linly-Talker-Stream](https://github.com/Kedreamix/Linly-Talker-Stream) | — | — | — | Linly-Talker 的**实时流式分仓**（WebRTC + barge-in 全双工），主仓停更后的迭代重心 |
| [HumanAIGC-Engineering/OpenAvatarChat-WebUI](https://github.com/HumanAIGC-Engineering/OpenAvatarChat-WebUI) | — | — | — | OAC 0.6.0 前后端分离后的前端仓库 |

---

## 16GB 可行性速查表

> 口径：`首帧延迟`/`FPS` 一律标注来源归属；`未查到` 表示无任何可验证来源。

| 仓库 | 类型 | License | Stars | **16GB 可行?** | 首帧延迟 | FPS | Streaming |
|---|---|---|---|---|---|---|---|
| [lipku/LiveTalking](https://github.com/lipku/LiveTalking) | 2D 视频驱动（wav2lip/musetalk/Ultralight 封装）+ WebRTC 推流 | **Apache-2.0** | 9.6k | ✅ **可行**。wav2lip 实测 1.3GB；musetalk **必须含 PR #612**（否则 17.6GB→OOM），修复后 3.9GB。⚠️ RAM 也是瓶颈（10s 1080p 缓存≈1.61GB/路，默认 5 路） | **未查到**（issue #368 的 6s~数分钟是 WebRTC 建链，非模型延迟） | wav2lip256: 3060=**60**、3080Ti=**120**；musetalk: 3080Ti=**42**、3090=**45**、4090=**72** `[作者实测]`；官方口径 inferfps/finalfps **≥25** 才算实时 | ✅ 支持**打断**+WebRTC/RTMP/虚拟摄像头+多并发（未见"双工"声明） |
| [duixcom/Duix-Avatar](https://github.com/duixcom/Duix-Avatar)（原 Duix.Heygem，官方改名） | 2D 视频驱动 + Docker 打包（**官方自述非实时合成**） | **DUIX.COM COMMUNITY LICENSE**（LICENSE 写 1000 MAU 门槛，README 写 10 万用户/$10M → **表述冲突，需法务确认**） | 16k / 15,547 | 🔴 **当前不可用**：默认 3 个镜像无 CC12.0 内核，[issue #624](https://github.com/duixcom/Duix-Avatar/issues/624)（2026-09-04）报 `no kernel image is available`，**Open 零回复**；另有 5090 专用 compose 但**未验证 5060 Ti**。且 16GB 单卡**无实测**；32GB RAM 官方标为「必要」 | **未查到** | **未查到**（[issue #251](https://github.com/duixcom/Duix-Avatar/issues/251)：4×24G 显卡"推理速度很慢"） | ❌ **官方定义非实时**；API 是提交+轮询；TTS `"streaming": false` |
| [anliyuan/Ultralight-Digital-Human](https://github.com/anliyuan/Ultralight-Digital-Human) | 2D 视频驱动，个性化轻量（移动端级） | ⚠️ **存疑**：shields 报 `not specified`，README 徽章写 Apache-2.0，LICENSE 文件抓取失败（HTTP 000） | 2.6k | ✅ **显存必然够**（**实测值未查到**；他人估算 ~1.5GB FP32）。真瓶颈是 **RAM ~10GB/路** 与软件栈 | **未查到** | <**10ms/帧**（2080 多并发，**需转 ONNX**）`[作者实测]`；FPS 数值未查到 | ⚠️ 支持但**代码未完善**；继任 **FeatherTalk** 已开源 C++ 流式；无打断 |
| [TMElyralab/MuseTalk](https://github.com/TMElyralab/MuseTalk) | 2D 视频驱动 latent inpainting 口型 | **MIT**（shields 报 `not identifiable`） | 6.6k | ✅ **可行且余量大**：**RTX 5060 Ti 16GB 直接实测 ~3.6GB**（361 帧验证）；官方 4GB fp16 最低；OpenTalking 峰值 5.0–5.2GB。**训练不可行**（≥32–85GB） | TTFV **1769ms**（3090）/ **2095ms**（4090）`[第三方 OpenTalking]`；原生未查到 | 官方 **30fps+@V100**；**5060 Ti 实测推算 ≈5.8fps**（7s音频→30s）；OpenTalking 3090=**28.87**、4090=**24.77**；瓶颈实测在 VAE decode 同步 | ❌ **官方成员确认不支持音频流式输入**（issue #33）；realtime 脚本仍事后合成音视频；无打断 |
| [bytedance/LatentSync](https://github.com/bytedance/LatentSync) | 2D 视频驱动 diffusion 口型 | **Apache-2.0** | 6.1k / 6,090 | ⚠️ **官方 1.6 门槛 18GB > 16GB**；但**第三方在 RTX 5070 Ti 16GB 实测跑通**：eager 峰值 **15,807MB**（贴 96%）、TensorRT 后 **12,551MB**。**1.5 官方门槛 8GB，稳妥可用**。训练不可行（20–55GB） | **未查到**（离线批处理，不适用） | 🔴 **≈0.61–0.90 帧/秒**，即 **27.7–41.1 秒换 1 秒成片（比实时慢 28~41 倍）**；GPU 利用率仅 1-3% | ❌ 无（离线批处理；长视频需自行切段 issue #326） |
| [Rudrabha/Wav2Lip](https://github.com/Rudrabha/Wav2Lip) | 2D 视频驱动 CNN 口型（2020 鼻祖） | 🔴 **无 LICENSE 文件；README 明确「any form of commercial use is strictly prohibited」** | 13k / 13,219 | `[推断]` 显存够（封装方实测 **1.3GB**）；🔴 **但官方栈 `torch==1.1.0` + `Python 3.6` 无法在 sm_120/Py3.12 运行**，等于重写。**官方未给任何显存数字** | **未查到** | 经 LiveTalking：3060=**60fps**；作者实测 4090=**750fps**@1.3GB。**本仓库无时延数字**（仅 issue #584 抱怨"slow and very GPU consuming"） | ❌ 自身不支持（整段批处理）；靠 LiveTalking 封装 |
| [KlingAIResearch/LivePortrait](https://github.com/KlingAIResearch/LivePortrait)（原 KwaiVGI，已改名） | **视频驱动**肖像动画（隐式 3D 关键点+warping，**非 lip-sync、非 3DGS**） | **MIT**，但 InsightFace 检测模型**仅限非商业研究**，商用须替换 | 19k / 19,084 | ✅ 显存余量充足。三条独立口径：官方权重 **499.6MB** / 第三方 pinokio **6GB** / aigcpanel 官方包标 **8G+** → **建议取 8GB 保守门槛** | **约 1 分钟** `torch.compile` 首次编译（官方）；业务首帧概念不存在 | **14.77ms/帧 ≈ 67.7 FPS**（4090+torch.compile，纯模型）`[官方实测]`；**实时改造版 FasterLivePortrait 在 3090 上 30+ FPS（含前后处理）** | ❌ 官方无；第三方 FasterLivePortrait(TensorRT 8.x)/FacePoke 有实时版 |
| [HumanAIGC-Engineering/OpenAvatarChat](https://github.com/HumanAIGC-Engineering/OpenAvatarChat) | 全链路实时对话框架（ASR+LLM+TTS+Avatar：LiteAvatar/LAM/MuseTalk/FlashHead） | **Apache-2.0** | 3.8k / 3,773 | ✅ **大多档可行**：**LAM 实测 3.1GB**、**LiteAvatar 实测 5.3GB**（RTX 3060 6G 笔记本实测）；配置档 LAM/Qwen-Omni **2–3GB**、本地TTS档 **8–10GB**、MuseTalk **~4GB**（另一口径 6-8GB）；❌ 仅 MiniCPM 全本地 **20GB+**（int4 10GB 可）。依赖要求 **N卡 4G+ / CUDA ≥12.4** | **平均响应 2.2s**（官方；测得于 i9-13900KF + RTX 4090 = 「用户说完→数字人开口」）；首帧未查到 | 输出典型 **25fps**（第三方） | ✅✅ **0.6.0 起全后端手动打断 + 双工打断**（本批最强） |
| [Kedreamix/Linly-Talker](https://github.com/Kedreamix/Linly-Talker) | 多后端对话数字人（SadTalker/Wav2Lip/MuseTalk/ER-NeRF） | **MIT** | 3.5k / 3,453 | ⚠️ **取决于组合，需实测**；官方无数字，第三方估算「**12GB 为最低门槛**」（LLM int4 6–8GB + TTS 2–4GB + ASR 1–2GB）`[第三方估算，可信度低]` | **未查到** | MuseTalk **V100 >30fps**（转述上游）；无自身实测 | ⚠️ 主仓 ❌；**[Linly-Talker-Stream](https://github.com/Kedreamix/Linly-Talker-Stream) ✅ barge-in 全双工** |
| [xszyou/Fay](https://github.com/xszyou/Fay) | agent 框架 + 数字人连接器（**不渲染**） | 🔴 **GPL-3.0**（README 自称「商用免责」但无许可证依据） | 14k / 13,545 | ✅ **本体 0 显存**；显存取决于外挂模型 | **未查到** | **未查到** | ✅ 全时流式 + 支持**打断** |
| [modstart-lib/aigcpanel](https://github.com/modstart-lib/aigcpanel) | Electron 桌面面板 + 本地模型启动包 / 云 API（集成 MuseTalk/LatentSync/Wav2Lip/Heygem/LivePortrait/SoulX-FlashHead 等 12 包） | **Apache-2.0**（不覆盖第三方模型） | 5.6k / 5,574 | ✅ **面板 0 显存**；**官方 12 个模型包中 11 个门槛 ≤16GB**（Wav2Lip/SadTalker 4G+、MuseTalk 6G+、LatentSync/Heygem-v2/LivePortrait/FlashHead-Lite 8G+、Heygem 16G+；唯一超限 FlashTalk-14B 24G+）；**6 个包标「支持50显卡」** | **未查到** | **未查到**（唯一 FPS 数字是 FlashHead Lite 标称 **96 FPS**） | ❌ 未查到（任务式批处理） |

### 速查表的四个「最」的结论
1. **最稳的 16GB 实时方案**：**lipku/LiveTalking（`--model wav2lip`）** —— 实测 1.3GB、3060 即 60fps、官方验证栈正是 CUDA 12.8（对齐 sm_120）、支持打断+中文+全身、2026-09 仍活跃。**用 musetalk 后端则必须确认已含 PR #612。**
2. **唯一有「同款硬件 + sm_120」直接实测的项目**：**MuseTalk** —— [issue #409](https://github.com/TMElyralab/MuseTalk/issues/409) 在 RTX 5060 Ti 16GB 上跑通 361 帧、**~3.6GB VRAM**；但吞吐 ≈**5.8fps，达不到实时**，**官方成员确认不支持音频流式**，且必须 **torch 2.10+cu128** 并绕开 mmcv（改 face-alignment）。
3. **显存最不可能成为瓶颈、但软件栈最容易翻车的**：**Ultralight-Digital-Human**（官方 cu117 栈不可用）、**Wav2Lip**（torch 1.1.0 + 严格非商用）。
4. **🔴 最需要警惕的「README 与现实的落差」**：**Duix.Avatar** —— README 明确写了 50 系部署方案与「verified on 5090」，但同型号 Blackwell(CC12.0) 的 [issue #624](https://github.com/duixcom/Duix-Avatar/issues/624) 直接报 `no kernel image is available` 且**零回复**；且官方自述**非实时**。（**且 aigcpanel 官方把 Heygem-v2 标为「支持50显卡」，与 #624 冲突 → 两处口径矛盾，必须实测。**）
5. **分发方官方门槛比上游 README 更可信、信息量更大**：**aigcpanel 的官方模型市场逐包列了 GPU 门槛**（12 包中 11 个 ≤16GB、6 包标「支持50显卡」），是目前**唯一按包给出 50 系可用性标注**的来源；这类"集成商实测/门槛表"值得优先采信。

---

## 16GB / sm_120 落地检查清单（给部署用）

| 检查项 | 要求 | 依据 |
|---|---|---|
| PyTorch | **≥2.7，且必须 cu128+ 轮子**（推荐 cu128–cu130） | [PyTorch 2.7 发布说明](https://pytorch.org/blog/pytorch-2.7/)；MuseTalk [#409](https://github.com/TMElyralab/MuseTalk/issues/409) 的 cu124/cu126 ❌ vs cu128 ✅；Duix [#624](https://github.com/duixcom/Duix-Avatar/issues/624) |
| onnxruntime-gpu | **≥1.27（CUDA 13 构建）**；避免 1.18/1.21 等旧档；或换第三方 sm_120 wheel | [ORT 官方版本表](https://onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider.html)；[Natfii/onnxruntime-gpu-blackwell](https://github.com/Natfii/onnxruntime-gpu-blackwell)；ORT [#27621](https://github.com/microsoft/onnxruntime/issues/27621) |
| CUDA | 驱动支持的 CUDA **≥12.8**（目标机 CUDA 13.2 满足） | OpenAvatarChat 快速开始页；Duix「50 series … CUDA 12.8」 |
| NVIDIA 驱动 | RTX 50 系建议 **≥575.64.03**（目标机 610.43.02 ✅） | [deepwiki: OAC System Requirements](https://deepwiki.com/HumanAIGC-Engineering/OpenAvatarChat/1.1-system-requirements) |
| flash-attn | 需**源码编译**（`--no-build-isolation`），sm_120 下耗时且易失败 | OpenAvatarChat 快速开始页 |
| mmcv / mmpose | **Python 3.12 无 wheel 且源码编译失败** → 改用 **face-alignment** | MuseTalk [#409](https://github.com/TMElyralab/MuseTalk/issues/409)、[#334](https://github.com/TMElyralab/MuseTalk/issues/334) |
| `torch.load` | 新 PyTorch 默认 `weights_only=True`，加载旧 checkpoint 需调整参数 | MuseTalk [#334](https://github.com/TMElyralab/MuseTalk/issues/334) |
| 逐帧显存 | 确认推理路径带 `@torch.no_grad()`，并把 `batch_size` 从 16 降到 8 | LiveTalking [PR #612](https://github.com/lipku/LiveTalking/pull/612) / [issue #192](https://github.com/lipku/LiveTalking/issues/192)；MuseTalk [PR #349](https://github.com/TMElyralab/MuseTalk/pull/349) |
| 验证方法 | **不要只信 README**：先在目标机复现 `python -c "import torch;print(torch.cuda.get_device_capability())"`，再跑官方最小推理样例 | Duix [#624](https://github.com/duixcom/Duix-Avatar/issues/624) 的教训 |
| 商用许可 | **Wav2Lip 明确非商用**；**Duix 有 1000 MAU / 10 万用户门槛（表述冲突）**；**LivePortrait 须替换 InsightFace 模型**；**Fay 是 GPL-3.0** | 各项目 LICENSE / README |
| 训练可行性 | **16GB 单卡一律不可训练**：MuseTalk ≥32GB、LatentSync ≥20GB、MuseTalk stage2 ≥54GB、LatentSync stage2_512 ≥55GB | 各项目 README |

---

## 来源清单

### A. GitHub 仓库与官方文件（逐字引用来源）
1. https://github.com/lipku/LiveTalking
2. https://raw.githubusercontent.com/lipku/LiveTalking/main/README.md
3. https://raw.githubusercontent.com/lipku/LiveTalking/main/README-EN.md
4. https://raw.githubusercontent.com/lipku/LiveTalking/main/requirements.txt
5. https://raw.githubusercontent.com/lipku/LiveTalking/main/LICENSE
6. https://doc.livetalking.ai/docs/faq/
7. https://github.com/duixcom/Duix-Avatar （原 `duixcom/Duix.Heygem`，301 重定向 + Release v1.0.5 改名声明确认）
8. https://raw.githubusercontent.com/duixcom/Duix.Heygem/main/LICENSE （DUIX.COM COMMUNITY LICENSE AGREEMENT）
9. https://github.com/duixcom/Duix.Avatar/releases
10. https://github.com/duixcom/Duix-Mobile ｜ https://raw.githubusercontent.com/duixcom/Duix-Mobile/main/README.md
11. https://github.com/duixcom/duix-skills
12. https://github.com/orgs/duixcom/repositories
13. https://github.com/anliyuan/Ultralight-Digital-Human （默认分支 `master`；`main` 不存在）
14. https://raw.githubusercontent.com/anliyuan/Ultralight-Digital-Human/master/README.md
15. https://github.com/anliyuan/FeatherTalk
16. https://github.com/TMElyralab/MuseTalk
17. https://raw.githubusercontent.com/TMElyralab/MuseTalk/main/README.md
18. https://raw.githubusercontent.com/TMElyralab/MuseTalk/main/requirements.txt
19. https://raw.githubusercontent.com/TMElyralab/MuseTalk/main/LICENSE （MIT）
20. https://huggingface.co/TMElyralab/MuseTalk
21. https://arxiv.org/abs/2410.10122 （MuseTalk 技术报告）
22. https://github.com/bytedance/LatentSync
23. https://raw.githubusercontent.com/bytedance/LatentSync/main/README.md
24. https://raw.githubusercontent.com/bytedance/LatentSync/main/requirements.txt
25. https://raw.githubusercontent.com/bytedance/LatentSync/main/LICENSE
26. https://huggingface.co/ByteDance/LatentSync-1.6
27. https://github.com/Rudrabha/Wav2Lip
28. https://raw.githubusercontent.com/Rudrabha/Wav2Lip/master/README.md
29. https://raw.githubusercontent.com/Rudrabha/Wav2Lip/master/requirements.txt
30. https://github.com/KlingAIResearch/LivePortrait （重定向自 `KwaiVGI/LivePortrait`）
31. https://raw.githubusercontent.com/KlingAIResearch/LivePortrait/main/readme.md （小写文件名；`README.md` 为 404）
32. https://raw.githubusercontent.com/KlingAIResearch/LivePortrait/main/requirements.txt
33. https://raw.githubusercontent.com/KlingAIResearch/LivePortrait/main/LICENSE
34. https://github.com/KlingAIResearch/LivePortrait/blob/main/assets/docs/speed.md
35. https://github.com/HumanAIGC-Engineering/OpenAvatarChat
36. https://raw.githubusercontent.com/HumanAIGC-Engineering/OpenAvatarChat/main/README.md
37. https://raw.githubusercontent.com/HumanAIGC-Engineering/OpenAvatarChat/main/readme_en.md
38. https://humanaigc-engineering.github.io/OpenAvatarChat/getting-started/
39. https://humanaigc-engineering.github.io/OpenAvatarChat/releases/release-notes
40. https://github.com/HumanAIGC-Engineering/OpenAvatarChat-WebUI
41. https://github.com/Kedreamix/Linly-Talker
42. https://raw.githubusercontent.com/Kedreamix/Linly-Talker/main/README.md
43. https://raw.githubusercontent.com/Kedreamix/Linly-Talker/main/LICENSE
44. https://raw.githubusercontent.com/Kedreamix/Linly-Talker-Stream/main/README.md
45. https://github.com/xszyou/Fay
46. https://raw.githubusercontent.com/xszyou/Fay/master/README.md
47. https://raw.githubusercontent.com/xszyou/Fay/master/LICENSE （GPL-3.0）
48. https://raw.githubusercontent.com/xszyou/fay-ue5/main/README.md
49. https://github.com/modstart-lib/aigcpanel
50. https://raw.githubusercontent.com/modstart-lib/aigcpanel/main/README.md
51. https://raw.githubusercontent.com/modstart-lib/aigcpanel/main/LICENSE
52. https://aigcpanel.com/zh/document/74 （官方：界面软件无硬件要求 / AI 模型需专业显卡 / 云模型无需本地显卡）
52b. 🔴 https://aigcpanel.com/zh/asset?tag=视频换口型 （**官方模型市场逐包 GPU 门槛表：12 包中 11 个 ≤16GB；6 包标「支持50显卡」；FlashHead Lite 96 FPS**）
52c. https://github.com/warmshao/FasterLivePortrait （**3090 上 TensorRT 30+ FPS（含前后处理）；TensorRT 必须 8.x；需自编译 grid_sample 插件与 onnxruntime**；默认分支 `master`）
52d. https://www.53ai.com/news/OpenSourceLLM/2025061631429 （**OpenAvatarChat 实测带图：RTX3060 6G 笔记本上 LAM 3.1GB / LiteAvatar 5.3GB；N卡 4G+ / CUDA ≥12.4**）

### B. GitHub Issues / PR（含实测数字，本报告核心证据）
53. 🔴 https://github.com/duixcom/Duix-Avatar/issues/624 （**RTX 5070 Blackwell CC12.0 `no kernel image is available`；2026-09-04 开，Open 零回复**）
54. https://github.com/duixcom/Duix-Avatar/issues/325 （6G VRAM can't run lite version）
55. https://github.com/duixcom/Duix-Avatar/issues/251 （4×24G 显卡"推理速度很慢"）
56. 🔴 https://github.com/lipku/LiveTalking/pull/612 （**musetalk 实时推理显存 17.6G→3.9G**，2026-08-20 merged）
57. https://github.com/lipku/LiveTalking/issues/192 （4090 musetalk OOM：17.77GiB allocated / 18.47GiB reserved）
58. https://github.com/lipku/LiveTalking/issues/368 （WebRTC 建链 6s~数分钟）
59. https://github.com/lipku/LiveTalking/issues/618 （中文 ASR/TTS 后端请求，2026 活跃）
60. 🔴 https://github.com/TMElyralab/MuseTalk/issues/409 （⭐ **RTX 5060 Ti 16GB / sm_120 实测：~3.6GB VRAM；7s 音频→30s 推理；361 帧端到端；cu124/cu126 ❌ vs cu128 ✅；mmcv Py3.12 失败**）
61. 🔴 https://github.com/TMElyralab/MuseTalk/issues/33 （**项目组成员 czk32611：「音频的流式处理我们暂时没有研究」；30fps vs 60fps 根因 = VAE decode GPU→CPU 同步**）
62. https://github.com/TMElyralab/MuseTalk/issues/392 （VRAM 需求；GTX 1060 3GB 可跑但极慢）
63. https://github.com/TMElyralab/MuseTalk/issues/310 （4090D 实时爆显存；实测可压到 11G；batch=15 约 20G；维护者确认音视频事后合成）
64. https://github.com/TMElyralab/MuseTalk/issues/334 （50 系支持；pytorch 2.7+cu128；mmcv 2.1.0；`torch.load` 参数）
65. https://github.com/TMElyralab/MuseTalk/pull/349 （推理加 `torch.no_grad()`）
66. https://github.com/TMElyralab/MuseTalk/issues/396 （社区质疑实时性）
67. https://github.com/TMElyralab/MuseTalk/issues/248 （4090 只能跑到 5fps）
68. https://github.com/TMElyralab/MuseTalk/issues/49 ｜ https://github.com/TMElyralab/MuseTalk/issues/333 ｜ https://github.com/TMElyralab/MuseTalk/issues/377 ｜ https://github.com/TMElyralab/MuseTalk/issues/359 ｜ https://github.com/TMElyralab/MuseTalk/issues/173
69. 🔴 https://github.com/bytedance/LatentSync/issues/365 （**RTX 5070 Ti 16GB 实测：eager 峰值 15,807MB / TensorRT 12,551MB；397.5s 换 9.68s 成片**）
70. https://github.com/bytedance/LatentSync/issues/335 （中文：「因為顯存不夠，只有16GB，想詢問是否能使用v1.5的版本？」）
71. https://github.com/bytedance/LatentSync/issues/314 ｜ https://github.com/bytedance/LatentSync/issues/284 ｜ https://github.com/bytedance/LatentSync/issues/278 ｜ https://github.com/bytedance/LatentSync/issues/326 ｜ https://github.com/bytedance/LatentSync/issues/270 ｜ https://github.com/bytedance/LatentSync/issues/137
72. https://github.com/Rudrabha/Wav2Lip/issues/584 ｜ https://github.com/Rudrabha/Wav2Lip/issues/613 （**追踪器事实停更：最新活动 2024-01 / 2023-11**）
73. https://github.com/anliyuan/Ultralight-Digital-Human/issues/120 （推理内存占用近 10G）
74. https://github.com/anliyuan/Ultralight-Digital-Human/issues/74 （推理全程走 CPU，GPU 无动静）
75. https://github.com/anliyuan/Ultralight-Digital-Human/issues/103 （催开源实时流代码）
76. https://github.com/anliyuan/Ultralight-Digital-Human/issues/101 ｜ https://github.com/anliyuan/Ultralight-Digital-Human/issues/18
77. https://github.com/microsoft/onnxruntime/issues/27621 （**CUDA EP 在 Blackwell RTX 5060 sm_120 上 silent deadlock**）
78. https://github.com/pytorch/pytorch/issues/174731 （RTX 5060 sm_120 预编译 cubin 问题）
78b. https://github.com/pytorch/pytorch/issues/166794 （**"[Bug] RTX 5070 Ti (sm_120) not recognized by PyTorch 2.5.1+cu121" —— 与 LatentSync 的 pin 完全同型**；仅索引标题级证据，正文页超时）
78c. https://github.com/Comfy-Org/ComfyUI/issues/7127 （cu121 运行期警告：支持的 arch 列表 **止于 sm_90，无 sm_120**；仅索引标题级证据）
78d. https://pypi.org/pypi/torch/1.1.0/json （**torch 1.1.0 仅 9 个 wheel、2019-04-30、仅 cp27–cp37、manylinux1**，证明 Wav2Lip 依赖栈已不可复现）

### C. 第三方基准 / 博客 / 文档（非官方）
79. https://www.cnblogs.com/livetalking/articles/22134438 （拆解 LiveTalking 资源消耗；三模型显存**估算**表；RAM 缓存 1.61GB/10s；NVENC 表；**作者 = lipku 本人**）
80. https://blog.csdn.net/lipku/article/details/148593974 （**作者本人 4090 实测**：wav2lip 1.3G/750fps、musetalk 12G/60fps；多线程共享模型）⚠️ 本次抓取被反爬（返回 2KB 壳页），内容引自搜索摘要与作者博客互证
81. https://datascale-ai.github.io/opentalking/latest/en/avatar_models/musetalk/ （MuseTalk 峰值显存 5.078/5.203GB；TTFV 1769/2095ms；Steady FPS 28.87/24.77）
82. https://datascale-ai.github.io/opentalking/latest/en/benchmark/
83. https://deepwiki.com/HumanAIGC-Engineering/OpenAvatarChat/1.1-system-requirements （**逐 Handler 显存表 + CUDA 12.8/RTX 50 系 + 驱动 575.64.03+ + 2.2s 延迟**）
84. https://deepwiki.com/HumanAIGC-Engineering/OpenAvatarChat/6-avatar-systems （LiteAvatar 3GB/会话；MuseTalk 6-8GB；LAM <1GB；floor(VRAM/3GB) 并发规则）
85. https://deepwiki.com/HumanAIGC-Engineering/OpenAvatarChat/2-getting-started （旧版 README 的配置档显存表）
86. https://deepwiki.com/HumanAIGC-Engineering/OpenAvatarChat/4.2-configuration-examples
87. https://deepwiki.com/anliyuan/Ultralight-Digital-Human/3.2-digital-human-model-training
88. https://raw.githubusercontent.com/ShmuelRonen/ComfyUI-LatentSyncWrapper/main/README.md （LatentSync 权重体积 ~5GB/~1.6GB；20GB VRAM 优化；FlashAttention-2 无 xFormers）
89. https://github.com/Natfii/onnxruntime-gpu-blackwell （官方 PyPI onnxruntime-gpu **不含 sm_120 内核**；1.24.1 prebuilt）
90. https://github.com/Lytanshade/LivePortrait-pinokio （第三方称 LivePortrait "6GB VRAM ~8GB install"）
91. https://www.compshare.cn/images/compshareImage-1cft3sk9gvta （Fay 官方镜像：40GB、最近更新 2025-12-05、CUDA 版本栏为空）
92. https://blog.csdn.net/weixin_42561464/article/details/156119020 （Linly-Talker 模块显存**估算**；「12GB 为最低门槛」；**AI 生成风格、无实测截图，可信度低**）
93. https://onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider.html （**ORT 版本→CUDA 版本对应表；1.27 起默认 CUDA 13.0**）
94. https://pytorch.org/blog/pytorch-2.7/ （**PyTorch 2.7 引入 Blackwell 支持 + CUDA 12.8 wheels + Triton 3.3**）
95. https://pypi.org/pypi/torch/json （torch 版本序列核对，2026-09 最新 2.14.0）
96. https://pypi.org/pypi/onnxruntime-gpu/json （onnxruntime-gpu 最新 1.30.0）
97. https://www.bilibili.com/video/BV1uUNEzwEjh/ （第三方标题：「OpenAvatarChat 一键包更新4 **新增支持50系** 新增 MuseTalk」；**页面正文未能抓取，仅作线索，未作为结论依据**）

### D. 统计口径来源
98. shields.io JSON 端点：`https://img.shields.io/github/stars/<owner>/<repo>.json`、`.../license/<owner>/<repo>.json`、`.../last-commit/<owner>/<repo>.json`（2026-09-22 抓取，原始返回值见各节 (a)）
99. GitHub commits Atom feed：`https://github.com/<owner>/<repo>/commits/<branch>.atom`（用于取精确到秒的最后提交时间）

### E. 明确「未查到」的项（未编造）
- **Ultralight-Digital-Human**：实测显存值、推理 FPS 数值、首帧延迟；**其 LICENSE 存在性存疑**（LICENSE 文件 HTTP 000，非 404）。
- **FeatherTalk**：显存/延迟数字（README 抓取返回 0 字节）。
- **LiveTalking**：原生首帧延迟；sm_120 专项报告（搜索 total=0，属「无已知 issue」而非「已适配」）。
- **MuseTalk**：原生脚本的首帧延迟（只有 OpenTalking 第三方口径）；sm_120 下的原生实时 FPS（只有 7s→30s 批量推算）。
- **Duix-Avatar/Heygem**：具体 VRAM GB 数、fp16 权重体积、延迟、FPS；Docker Hub 镜像体积；**LICENSE 正文与 README 商用阈值的对应关系**。
- **LatentSync**：首帧延迟；是否存在 1.7 版本。
- **Wav2Lip**：本仓库实测显存、时延具体数字。
- **LivePortrait**：官方显存要求与官方 nvidia-smi 实测、业务首帧延迟（现有门槛来自分发方/第三方：pinokio 6GB、aigcpanel 包标 8G+）。
- **OpenAvatarChat**：MuseTalk 后端确切 GB（存在 4GB 与 6-8GB 两个第三方口径）、首帧延迟、全身支持、sm_120 专项 issue。
- **Linly-Talker**：官方训练/推理显存、延迟、首帧。
- **Fay**：官方任何延迟/FPS 数字；xuniren/metahuman 连接器名称级原文。
- **aigcpanel**：各集成模型的具体显存 GB、延迟、FPS、流式能力。
- **方法论局限**：GitHub REST API 未认证额度（60/hr）在本次会话中途耗尽，**未能对全部仓库执行 `search/issues` 关键词检索**（`sm_120`/`Blackwell`/`OOM`/`16GB`）；`duixcom` org 完整仓列表未能拉全（改用逐仓探测，**可能遗漏未探测的仓**）。LivePortrait、Linly-Talker、Fay、aigcpanel 的 sm_120 判断部分依赖依赖清单与搜索片段推断，**建议后续用带 token 的 API 补做 issue 检索验证**。
- **未采信来源**：gitcode「MuseTalk 4090 性能优化」页面**自述由 AIGC 生成**且需登录看全文 → 已弃用。
