# 实时数字人开源项目 · 16GB 显卡选型调研（C1 分片）

**目标硬件**：RTX 5060 Ti 16GB（sm_120 / Blackwell）、32GB RAM、Ubuntu、NVIDIA 驱动 610.43.02、CUDA 13.2
**调研日期**：2026-09-22　**责任人**：C1　**范围**：仅 3 个仓库（lipku/LiveTalking、anliyuan/Ultralight-Digital-Human、TMElyralab/MuseTalk）

**标注约定**：`[作者自述]` = 项目作者/官方文档；`[第三方实测]` = 用户在 issue/博客的实测；`[推算]` = 由可验证数字推导，非直接引用；`未查到` = 未找到可验证来源。

---

## 0. 2026 年状态核实（最后提交时间 / Star / License）

测量方式：`curl -sL "https://github.com/<owner>/<repo>/commits/<branch>.atom"` 取 `<updated>` 首条（提交时间戳，**无 API rate limit**）；Star/License 取 shields.io `img.shields.io/github/{stars,license,last-commit}/<owner>/<repo>.json`。核实日期：**2026-09-22**。

| 仓库 | 默认分支 | **最后提交时间** | Star | License |
|---|---|---|---|---|
| **lipku/LiveTalking** | `main`（`master` 不存在） | **2026-09-13T11:56:50Z**（距调研日仅 9 天） | **9.6k** | Apache-2.0 |
| **anliyuan/Ultralight-Digital-Human** | **`master`**（`main` 不存在） | **2026-07-22T12:29:23Z** | **2.6k** | shields.io 返回 **not specified**（README 徽章指向 `./LICENSE`，但本次 `curl` 该文件返回 `000`/连接失败而非 404，**许可证状态存疑，未验证**；对比 LiveTalking 的 `LICENSE` 返回 `200` + Apache License 正文） |
| **TMElyralab/MuseTalk** | `main`（`master` 不存在） | **2025-09-26T05:44:17Z**（⚠️ **已近一年未更新**） | **6.6k** | shields.io 返回 **not identifiable by github**（HF 页标注 `creativeml-openrail-m`） |
| anliyuan/FeatherTalk（Ultralight 继任者，参考项） | — | 2026-09（shields "september"） | **93** | Apache-2.0 |

**结论 / 对选型的影响**：
1. **LiveTalking 是三者中唯一在 2026 年持续高频维护的项目**（最后提交 2026-09-13，且 PR #612 于 2026-08-20 合并）。目标硬件 sm_120 属于新平台，**维护活跃度是能否踩到坑并被修好的关键变量**，此项 LiveTalking 明显占优。
2. **MuseTalk 上游代码已停滞近一年**（2025-09-26 之后无提交）。其 2026 年的 sm_120 适配进展**全部来自社区 issue/PR 自述**（如 #409、#334），**不存在官方修复承诺**。选它需接受"自己打补丁"的现实。
3. **Ultralight-Digital-Human 在 2026 年仍有更新**（2026-07-22），但作者已把重心转向 FeatherTalk；本仓库的「流式代码未完善」缺口在其继任者中才补齐。
4. **默认分支陷阱**：Ultralight 默认分支是 **`master`**，其 `main` 分支不存在（`README.md` 请求返回 404，仅 14 字节）。自动化脚本若硬编码 `main` 会拿到 404 —— 本报告初次抓取即撞上此坑。

---

## 1. lipku/LiveTalking

> 定位：**实时交互流式数字人引擎**（不是单一模型，而是 wav2lip / musetalk / Ultralight 三模型的流式封装 + WebRTC/RTMP 推流管线）。

### (c) 显存需求 + 16GB 单卡能否跑

**[作者自述] 官方性能表**（README「6. 性能指标」）：需要区分「模型本体显存」与「推理峰值显存」。作者本人（博客园 ID「恒中」= lipku）在博客中给出了**按源码估算**的三模型显存：

| 模型 | 技术路线 | 显存占用（估算） | 精度 | 输入分辨率 |
|---|---|---|---|---|
| Wav2Lip | CNN 编解码 + 空间注意力 | **~200MB** | FP32 | 256x256 |
| MuseTalk | UNet (Diffusion) + VAE + Whisper | **~1.2GB（FP32）/ ~600MB（FP16）** | FP16 | 256×256 |
| UltraLight | 轻量 UNet + HuBERT | **~1.5GB** | FP32 | 160×160 |

原文逐字引用：「MuseTalk | UNet (Diffusion) + VAE + Whisper | **~1.2GB（FP32）/ ~600MB（FP16）** | FP16 | 256×256」
来源：[拆解 LiveTalking 的资源消耗：GPU、CPU、内存、带宽到底花在哪？](https://www.cnblogs.com/livetalking/articles/22134438)
⚠️ **注意**：该表标注为「显存占用（估算）」，是作者读源码的估算值，**不是实测峰值**。

**[第三方实测] musetalk 真实峰值显存 —— 17.62 GB → 3.9 GB**（这条最关键）：
PR #612 标题逐字：「fix: MuseReal.inference_batch 缺 @torch.no_grad()，实时推理显存 **17.6G→3.9G**」，已于 **2026-08-20 被作者 lipku 合并（MERGED）**。PR 正文逐字引用：

> 「问题：同文件的 warm_up / __init__ 都带 no_grad，唯独实时推理路的 inference_batch 漏了，调用方 BaseAvatar.inference 也没有 ⇒ 每批推理都在构建 autograd 图，显存被求导图白白占掉。
> **实测（RTX 5090，musetalk，batch_size=16）：torch allocated 峰值 17.62 GB → …**」

来源：[lipku/LiveTalking PR #612](https://github.com/lipku/LiveTalking/pull/612)（第四方摘录自 PR 页 `twitter:description` 元数据，正文尾部被页面截断，但两个数字完整可读）

**[第三方实测] 未修复时 4090 直接 OOM**（issue #192，2024-08-07）：
> 「RuntimeError: CUDA out of memory. Tried to allocate 256.00 MiB (GPU 0; 23.65 GiB total capacity; **17.77 GiB already allocated**; 45.19 MiB free; **18.47 GiB reserved** in total by PyTorch)」

社区给出的解法（issue #192 评论，leos-code）逐字：
> 「在 musereal.py 的 inference方法上添加 @torch.no_grad() 注解；启动app.py时，设置batch_size为8或更小，我设置8可以正常运行」

来源：[LiveTalking issue #192「musetalk 启动后 oom」](https://github.com/lipku/LiveTalking/issues/192)

**[作者自述] 内存（RAM）与并发上限 —— 与显存同等重要的是内存**（来源同上[博客园文章](https://www.cnblogs.com/livetalking/articles/22134438)）：
作者从源码逐字拆解：系统启动时会把整个 avatar 视频所有帧加载进内存。以 **10 秒 1080p 视频（25fps）** 为例：
> 「总帧数：10s × 25fps = 250 帧
> 全帧缓存：250 × 1920 × 1080 × 3 字节 = 250 × 6.2 MB ≈ **1.56 GB**
> 人脸缓存：250 × 256 × 256 × 3 字节 = 250 × 0.2 MB ≈ **49 MB**
> 坐标数据：250 × 4 × 4 字节 ≈ 4 KB（可忽略）
> ─────────────────────────────────────────────
> 总计：**约 1.61 GB（仅一个avatar视频）**」

并发默认值逐字：「5 路（默认 max_session）」。→ **32GB RAM 的机器上，多路数字人时 avatar 帧缓存是主要内存开销（每路约 1.6GB/10 秒 1080p 素材），显存反而不是瓶颈。**

**NVENC 硬编并发上限（若把 libx264 软编换成 `h264_nvenc`）**（同来源）：

| GPU | NVENC 芯片数 | 并发编码上限 | 说明 |
|---|---|---|---|
| RTX 3060 | 1 个 | 3 路 | 消费卡驱动限制 |
| RTX 3080Ti / 3090 | 1 个 | 3 路 | 消费卡驱动限制 |
| RTX 4090 | 2 个 | 5-8 路 | 双编码器 + 新驱动放宽 |
| RTX A5000 / A6000 | 1-2 个 | 20+ 路 | 专业卡无驱动限制 |
| A100 / H100 | 无 NVENC | 0 路 | 纯计算卡，没有编码器！ |

⚠️ 关键提醒（逐字）：「A100/H100 是纯 CUDA 计算卡，**没有 NVENC 编码器**。如果选用 A100 做高并发推理，H.264 编码依然要靠 CPU 软编。」
→ **RTX 5060 Ti 未在该表中，按消费卡同类推断为 1 个 NVENC / 3 路并发上限**（[推算]，表中无 50 系数据，属未验证项）。默认走 `libx264` 纯软件编码，作者测算**单路 1080p 推流持续占用约 0.45 个 x86 CPU 核心**，5 路约 2.0-2.5 核。

**[作者自述] 反压机制**：输出缓冲区积压超过 5 帧时主动休眠（`base_avatar.py:482-484`）：`if buffer_size >= 5: time.sleep(0.04 * buffer_size * 0.8)`。每会话 4 线程（`render()` / `inference()` / `process_frames()` / `process_tts()`）。

**16GB 单卡结论**：
- wav2lip256 / Ultralight：**毫无压力**（作者估算 ~200MB / ~1.5GB）。
- musetalk：**能跑，但有前提**。修复前的默认配置峰值 17.6–18.5GB，**16GB 卡会 OOM**；`@torch.no_grad()` 修复已于 2026-08-20 合并进 main，修复后峰值 **3.9GB**，16GB 富余充足。若使用旧版本，需手动加 `@torch.no_grad()` 并把 `batch_size` 降到 8 或更低（issue #192 实测 batch_size=8 可正常）。作者博客亦指出「batch_size 默认 16，显存占用 ≈ O(batch_size)」。
- 另有一条社区说法「[4090运行不了的，得V100](https://github.com/lipku/LiveTalking/issues/192)」——该说法与官方 FPS 表（4090 musetalk 72fps 实时）矛盾，属早期未打 no_grad 补丁时的现象，**不应采信**。

### (d) 端到端延迟：首帧延迟 + 推理 FPS

**[作者自述] 官方 FPS 表**（README「6. 性能指标」，逐字）：

| 模型 | 显卡 | FPS |
|:------|:------|:----|
| wav2lip256 | RTX 3060 | **60** |
| wav2lip256 | RTX 3080Ti | **120** |
| musetalk | RTX 3080Ti | **42** |
| musetalk | RTX 3090 | **45** |
| musetalk | RTX 4090 | **72** |

官方补充逐字：「后端日志 `inferfps` = GPU 推理帧率, `finalfps` = 最终推流帧率，**两者均需 >=25 才算实时**」。
推荐配置逐字：「wav2lip256 推荐 RTX 3060 及以上；musetalk 推荐 RTX 3080Ti 及以上」。

**[作者自述] 每帧 40ms 预算**：「每一帧只有 **40ms** 的预算（25fps）。这个流水线上的任何一个环节超时，用户就会感知到卡顿。」（同上博客园文章）

**首帧延迟（first-frame latency）：未查到官方数字**。
最接近的第三方现象记录是 issue #368，用户描述 WebRTC 首画建立耗时波动极大：
> 「webrtc从点击连接到前端有画面显示的过程中，有时候超级慢要几分钟，有时候只要6s的样子」

来源：[LiveTalking issue #368](https://github.com/lipku/LiveTalking/issues/368)（**注意：6s 是 WebRTC 链路建立耗时，不是模型首帧推理耗时**，不可当模型延迟引用）

⚠️ 官方 FPS 表**未标注测试分辨率**，也未覆盖 sm_120 显卡（无 50 系数据）。

### (e) 流式 / 可打断

**[作者自述] 支持，且是核心卖点**（README「Features」逐字）：
- 「2. 支持声音克隆」
- 「**3. 支持数字人说话被打断**」
- 「5. 支持 WebRTC、RTMP、虚拟摄像头输出」
- 「7. 支持多并发」

「支持打断重说」在「AI 数字人客服」场景中再次明确。架构上 `/human`（文本，echo/chat 模式）、`/humanaudio`（音频文件直接播放）接口，每个连接分配唯一 `sessionid` 支持多用户并发。

### (f) 中文支持 / 半身-全身

- **中文：完整支持**。README 为中文主文档；TTS 引擎模块化，支持 EdgeTTS、GPT-SoVITS、CosyVoice、腾讯云等中文方案；LLM 引擎对接 Qwen。2026 年仍有中文 issue 在活跃讨论（如 [issue #618](https://github.com/lipku/LiveTalking/issues/618) 请求接入 audio.cpp 作为本地 ASR + TTS / 声音克隆后端）。
- **全身：明确支持**。Features 第 4 条逐字：「**4. 支持全身视频拼接**」。
- 半身：通过自定义数字人形象（`avatar.html` 上传视频自动生成形象）实现，README 未单独列出「半身」词条。

### (g) 已知坑：sm_120 / CUDA / torch / flash-attn / onnx

- **[作者自述] 官方验证环境**（README「1. 安装」逐字）：「已在 **Ubuntu 22.04、Python 3.12、PyTorch 2.9.1、CUDA 12.8** 测试通过」，安装命令为 `pip install torch==2.9.1 torchvision==0.24.1 torchaudio==2.9.1 --index-url https://download.pytorch.org/whl/cu128`。**cu128 是 sm_120 可用的最低 CUDA 线**，这与目标机 CUDA 13.2 / 驱动 610.43.02 兼容（向下兼容）。
- **sm_120 专项 issue：未查到**。GitHub 搜索 `repo:lipku/LiveTalking sm_120` 返回 **total 0**；搜索 `Blackwell` 亦无结果。说明该仓库尚无 Blackwell 相关已知问题记录，属于「无明确阻塞报告」而非「已明确适配」。
- **依赖极轻，无 flash-attn / triton / onnx / tensorrt 依赖**。`requirements.txt` 内容为 python-dotenv、pyyaml、numpy、tqdm、scipy、resampy、soundfile、einops、ffmpeg-python、torchaudio 等（**注意：requirements.txt 中 torch 相关行被注释，torch 由用户按 CUDA 版本单独安装**）。`face_alignment`、`ninja`、`pyaudio` 等亦为注释状态。→ **不存在 flash-attn/triton 编译坑**。
- **已知维护性坑**：musetalk 路径的 `@torch.no_grad()` 缺失是长期坑（issue #192 于 2024-08 报告，PR #612 直到 **2026-08-20** 才合并）→ **务必确认所用版本已包含 PR #612**。
- 2026 年仍在活跃维护（PR #612 于 2026-08-20 合并入 main），这是三仓库中维护活跃度证据最明确的一个。
- 官方 FAQ 独立站：[doc.livetalking.ai/docs/faq](https://doc.livetalking.ai/docs/faq/)。

---

## 2. anliyuan/Ultralight-Digital-Human

> 定位：**超轻量级数字人**，主打移动端实时。README 明确其为「personal」路线（每人一段视频训一个专属模型）。作者已开新项目 **FeatherTalk**（C++ 流式推理）作为继任者。
> ⚠️ **默认分支为 `master`**（`main` 分支下 README 返回 404）。

### (c) 显存需求 + 16GB 单卡能否跑

**显存（VRAM）实测数字：未查到**。仓库 README、作者文档均**未给出任何显存数字**；也未找到第三方实测报告。

可用的周边证据：
- **[作者自述] 模型体积极小**（README「Updates」逐字）：「我重新设计了音频编码器与图像编解码结构，**audio + UNet 两个模型的总体积控制在 1M 以内**，推理更快」（指即将开源的超轻量流式版本）。FeatherTalk 的 **FeatherHuBERT** 相较 HuBERT Large「音频侧计算量约降低 **40 倍**」。
- **[作者自述] 显存无关的旁证（在 LiveTalking 中的估算）**：LiveTalking 作者把 UltraLight 模型显存估为 **~1.5GB（FP32）**，输入分辨率 160×160（见上文[博客园文章](https://www.cnblogs.com/livetalking/articles/22134438)）。**这是 LiveTalking 作者的估算，非本仓库自述，仅供参考**。
- **[第三方实测] 内存（RAM）而非显存才是坑**（issue #120，逐字）：
  > 「有没有发现内存占用巨高」「1分半的1080p视频训练完推理，**内存占用近10G**，虽然gpu算力要明显下降可是这内存占用显得很不ultra light啊」

  来源：[Ultralight-Digital-Human issue #120](https://github.com/anliyuan/Ultralight-Digital-Human/issues/120)
  → 32GB RAM 的机器单路推理 ~10GB 内存，**可用但需注意并发**。
- **[作者自述] 训练成本**：`train.py` 默认 `--batchsize 1`、`--epochs 200`（[DeepWiki: Training](https://deepwiki.com/anliyuan/Ultralight-Digital-Human/3.2-digital-human-model-training)）。默认 batch=1 说明训练端显存需求极低。**训练显存具体数字未查到**。

**16GB 单卡结论**：**显存侧必然够用**（模型是移动端级别），但**没有可引用的实测显存数字**，报告中应标注「未查到实测值，推断极小」。真正的限制在**软件栈兼容性**（见 (g)）与**内存占用**。

### (d) 端到端延迟：首帧延迟 + 推理 FPS

- **[作者自述] 每帧 <10ms（关键实时性证据）**（README「关于流式推理」，逐字）：
  > 「因为一般用到流式推理的场景一般对实时性要求比较高，所以这里我只写了wenet作为音频编码器的情况（**实测在2080这样的机器上多个并发时每帧音频处理+视频处理耗时10ms以内**，需要将模型转为onnx）」

  → 10ms/帧 < 40ms（25fps）预算，**实时性有余量**；且是「多个并发」下的数字。**注意前提：需将模型转为 ONNX**。
- **[作者自述] 推理效果/速度取向**：「In my experiments, hubert performs better, but wenet is faster and can run in real time on mobile devices.」
- **首帧延迟（first-frame latency）：未查到**。
- **推理 FPS 具体数值：未查到**（README 只有「10ms/帧」这一时间口径）。

⚠️ **重大坑（会直接毁掉实时性体验）**（issue #74 逐字）：
> 「推理时，使用文档中的语句：`python inference.py ...` 但是**全部用的是CPU去处理，GPU没动静**，是需要配置参数吗」

来源：[issue #74「推理时能用GPU吗」](https://github.com/anliyuan/Ultralight-Digital-Human/issues/74)
→ 官方 `inference.py` 默认**跑在 CPU 上**，需自行改 device。

### (e) 流式 / 可打断

- **[作者自述] 支持流式，但代码不完善**。README 逐字：「**这个模型是支持流式推理的，但是代码还没有完善**，之后我会提上来。」（该说明长期存在）
- 流式使用的实践建议（README 逐字）：「使用流式推理时，建议把静音的图片和对应的关键点放在单独的目录里，img_inference和lms_inference里。」
- **[第三方] 社区仍在催流式代码**：[issue #103](https://github.com/anliyuan/Ultralight-Digital-Human/issues/103) 标题逐字「还没有开源实时流吗？在不开源我要开源我的自己写的接的实时流的了」。
- **继任项目 FeatherTalk 已开源 C++ 流式推理代码**（README 首行逐字）：「**c++流式推理代码已开源到[FeatherTalk]**🎉…If you want a newer and more deployment-friendly version of this project, please check out FeatherTalk.」
- **可打断（barge-in）：未查到**。本仓库无打断机制描述；LiveTalking 才提供「说话被打断」。

### (f) 中文支持 / 半身-全身

- **中文：完整支持**（天然）。README 主文档为中文；官方 demo 视频使用「康辉老师的口播」（中文）；`data_utils/process.py` 走 wenet/hubert 中文友好链路。
- **半身 / 全身：未查到明确支持**。README 的训练数据要求逐字：「必须保证视频中**每一帧都有整张脸露出来**的人物」——即**面部/口播半身场景**。开源推理脚本没有全身拼接能力（对比 LiveTalking 有「支持全身视频拼接」）。
- 效果警告（README 逐字）：「如果你视频中声音质量比较差的话，效果大概率不会好」——需外接麦克风录制训练素材。

### (g) 已知坑：sm_120 / CUDA / torch / flash-attn / onnx

- **🔴 sm_120/Blackwell 兼容性：高风险，官方环境原样不可用**。README 指定的环境逐字：
  ```bash
  conda create -n dh python=3.10
  conda install pytorch==1.13.1 torchvision==0.14.1 torchaudio==0.13.1 pytorch-cuda=11.7 -c pytorch -c nvidia
  conda install mkl=2024.0
  pip install numpy==1.23.5
  ```
  **PyTorch 1.13.1 + CUDA 11.7 不含 sm_120 内核**，在 RTX 5060 Ti 上会直接报 `no kernel image is available`。作者逐字：「我是在1.13.1版本的pytorch跑的，其他版本的pytorch应该也可以」——**「应该也可以」是作者猜测，非验证结论，需自行迁移到 cu128 栈并重测**。
- **Python 3.10 锁定**（README 徽章与 env 创建命令）。numpy 锁 1.23.5，与新版 torch 常有冲突。
- **无 flash-attn / triton 依赖** → 无相关编译坑。
- **ONNX 是流式实时推理的前置条件**：作者明确「需要将模型转为onnx」才能达到 2080 上 <10ms/帧；README 安装列表含 `pip install onnxruntime`。用 wenet 编码器时必须**视频帧率 20fps**，用 hubert 必须 **25fps**，否则特征对齐错误。
- **SyncNet 已被作者移除**（README 逐字）：「I removed SyncNet-related code and docs to keep the repo simpler.」→ 若参照旧教程加 `--use_syncnet` 会失败（DeepWiki 仍记载该参数，**已过时**）。
- 训练时长/质量坑（[issue #18](https://github.com/anliyuan/Ultralight-Digital-Human/issues/18) 逐字）：「**200epochs是不是不太够**，训练了3个都有点抖」。
- 推理端音频特征提取慢（[issue #101](https://github.com/anliyuan/Ultralight-Digital-Human/issues/101) 逐字）：「感觉推理部分做音频特征提取速度太慢了，这部分如何改才能实现实时，有接入流式音频的可能性吗」；另有 [issue #161](https://github.com/anliyuan/Ultralight-Digital-Human/issues/161) 比较 Whisper-tiny 与 WeNet 速度（Whisper-tiny 快 2 倍）。

---

## 3. TMElyralab/MuseTalk

> 定位：**音频驱动口型同步（lip-sync）模型**，在 `ft-mse-vae` latent 空间做单步 inpainting，**不是 diffusion 模型**。被 LiveTalking 作为三种可选后端之一封装。
> ⭐ **本项目有 RTX 5060 Ti 16GB / sm_120 的直接实测报告（issue #409），是本报告最有价值的证据。**

### (c) 显存需求 + 16GB 单卡能否跑

**① [第四方实测 · 目标硬件直接命中] RTX 5060 Ti 16GB 上 ~3.6GB VRAM**
MuseTalk issue #409 标题逐字：「**Field Report: MuseTalk V1.5 working on RTX 5060 Ti (Blackwell sm_120) with Python 3.12 + mediapipe patch**」（2026-03-18，作者 chefboyrdave21）。正文「Performance」段落逐字：

> 「- **7sec audio → 30sec inference → MP4 output**
> - **~3.6GB VRAM for MuseTalk models**
> - Reference image face detection cached after first call」

文末环境逐字：「Setup: **RTX 5060 Ti 16GB | Ubuntu 24.04 | Python 3.12.3 | PyTorch 2.10.0+cu128 | MuseTalk V1.5**」

来源：[TMElyralab/MuseTalk issue #409](https://github.com/TMElyralab/MuseTalk/issues/409)

**② [同一 issue 内 2026-08-05 追加实测]** 用户 HeimdallCore 在同 issue 评论（2026-08-05）逐字：
> 「We ran into the same Blackwell (RTX 5060 Ti, sm_120) situation... **We verified this end-to-end (real MuseTalk V1.5 inference, 361 frames, correct lip-sync in the output).**」
→ 361 帧端到端跑通，进一步确认 16GB sm_120 可跑。

**③ [官方自述] 最低 4GB VRAM（fp16）**
README「Gradio Demo」段落逐字：
> 「For minimum hardware requirements, we tested the system on a Windows environment using an **NVIDIA GeForce RTX 3050 Ti Laptop GPU with 4GB VRAM**. In **fp16 mode, generating an 8-second video takes approximately 5 minutes**.」
来源：[MuseTalk README](https://github.com/TMElyralab/MuseTalk#gradio-demo)（下载页：[README 原文](https://raw.githubusercontent.com/TMElyralab/MuseTalk/main/README.md)）

**④ [第三方实测] 3GB 老卡也能跑（极慢）**
issue #392 逐字：
> 「I have only a gtx 1060 3gb PC. Is that enough to run this model?」
> 「**I was able to run musetalk with a gtx 1060 3gb. It's just ultra slow.**」
来源：[MuseTalk issue #392「What are the VRAM requirements」](https://github.com/TMElyralab/MuseTalk/issues/392)

**⑤ [第三方实测] 实时推理 11GB 足够，瓶颈是算力而非显存**
issue #310 用户 wanlichina 逐字：
> 「4090D的运算能力，你哪怕改成4或者2都能满足你的试试推理要求。**显存占用我测试下来，能做到11G。也就是3080,4080,5080都能跑，但是瓶颈不在显存，在GPU算力**，因为GPU占用率爆了，多开满足不了实时推理性能」
同 issue 楼主 codestart-zhu 未调 batch 时逐字：`torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 400.00 MiB (GPU 0; 23.64 GiB total capacity; 22.15 GiB already allocated...)`；调到 batch_size=15 后「现在占用在20g左右」。
来源：[MuseTalk issue #310「实时推理4090d爆显存问题」](https://github.com/TMElyralab/MuseTalk/issues/310)

**⑥ [推算] fp16 权重体积 —— 由 HuggingFace 上文件实际字节数推导**

| 文件 | 实测字节数 | 换算 | fp16 推算 |
|---|---|---|---|
| `musetalkV15/unet.pth` | **3,400,074,924 B** | 3.40 GB / 3.17 GiB | **≈1.70 GB（fp16）** |
| `musetalk/pytorch_model.bin`（v1.0） | **3,400,076,549 B** | 3.40 GB / 3.17 GiB | ≈1.70 GB（fp16） |

测量方式：对 `https://hf-mirror.com/TMElyralab/MuseTalk/resolve/main/...` 发 HTTP HEAD 读取 `Content-Length`。
→ 3.4GB ≈ SD1.4 UNet 的 **FP32** 体积（~860M 参数 × 4B）。因此**发布的 unet.pth 是 fp32 权重**，fp16 约为其一半 ≈ **1.7GB**。此为**推算值**，非官方标注。
（HF 官方仓库文件清单见 [huggingface.co/TMElyralab/MuseTalk](https://huggingface.co/TMElyralab/MuseTalk)，文件仅为 `musetalkV15/unet.pth`、`musetalkV15/musetalk.json`、`musetalk/pytorch_model.bin`、`musetalk/musetalk.json`。）

**⑦ [官方自述] 训练显存 —— 16GB 单卡完全不可能训练**
README「GPU Memory Requirements」逐字（基于 8× NVIDIA H20 测试）：

| Stage 1 | Batch Size | Grad Accum | Memory per GPU |
|---|---|---|---|
| | 8 | 1 | **~32GB** |
| | 16 | 1 | **~45GB** |
| | 32 | 1 | **~74GB** ✓ |

| Stage 2 | Batch Size | Grad Accum | Memory per GPU |
|---|---|---|---|
| | 1 | 8 | **~54GB** |
| | 2 | 2 | **~80GB** |
| | 2 | 8 | **~85GB** ✓ |

**16GB 单卡结论**：
- **推理：完全够用，且富余很大。** 目标硬件的直接实测为 **~3.6GB**（issue #409）；OpenTalking 第三方基准峰值 **5.078–5.203GB**；作者称 fp16 最低 4GB 可跑；极端情况（batch 过大）才会冲到 11–22GB。**16GB 单卡推理无压力**。
- **训练：单卡 16GB 不可行。** 官方最低档 stage1 bs=8 需 ~32GB，stage2 需 ~54GB 起。**16GB 需放弃自训练，只能用官方权重**（推理不需要训练）。

### (d) 端到端延迟：首帧延迟 + 推理 FPS

**① [官方自述] 30fps+ on Tesla V100** —— README 出现两次，逐字：
- 开头：「We introduce MuseTalk, a **real-time high quality** lip-syncing model (**30fps+ on an NVIDIA Tesla V100**).」
- Overview 第 3 条：「supports real-time inference with **30fps+ on an NVIDIA Tesla V100**.」
- Real-time Inference 注意事项第 3 条：「The generation process can achieve **30fps+ on an NVIDIA Tesla V100**」
⚠️ 官方口径**只给了 V100**，未给 3060/4090/sm_120 数据。

**② [第三方实测 · 目标硬件] RTX 5060 Ti 16GB：约 5.8 fps（非实时）**
由 issue #409 逐字数据推算：`7sec audio → 30sec inference`。音频 7 秒 @25fps = 175 帧，30 秒推理 → **175/30 ≈ 5.8 fps**（[推算]，基于第三方实测原值）。**这是本报告中唯一的目标显卡直接可算的吞吐数据，且结论是「远未达到实时」**。
⚠️ 需注意该方法用的是批量推理而非 `realtime_inference.py` 流式脚本，且原文未说明是否启用 fp16/优化。

**③ [第三方实测] LiveTalking 封装下的 musetalk FPS**（见上文 LiveTalking 官方表）：RTX 3080Ti **42** / RTX 3090 **45** / RTX 4090 **72**。

**④ [第三方 ] 首帧延迟（TTFV）实测 —— 这是本项目唯一可引用的首帧数字**
[OpenTalking 官方 benchmark](https://datascale-ai.github.io/opentalking/latest/en/benchmark/) 的 MuseTalk 页表格（该页说明逐字：「The numbers below are summarized from Benchmark. Steady FPS is model-generation throughput, not full user-perceived latency; STT, LLM, TTS, queueing, and WebRTC still affect the complete experience.」）：

| Hardware | Backend | Output | Steady FPS | First-turn total/ms | **TTFV/ms** | **Peak inference VRAM/GB** |
|---|---|---|---|---|---|---|
| RTX 3090 | OmniRT | 512×512 / 25fps | **28.868** | 3235.518 | **1769.484** | **5.078** |
| RTX 4090 | OmniRT | 512×512 / 25fps | **24.767** | 3605.564 | **2095.522** | **5.203** |
| NPU 910B2 | OmniRT | 512×512 / 25fps | **12.276** | 5781.453 | **4211.721** | **8.754** |

逐字引用：「Peak inference VRAM/GB | **5.078**」「TTFV/ms | **1769.484**」「First-turn total/ms | **3235.518**」
来源：[OpenTalking · MuseTalk 文档](https://datascale-ai.github.io/opentalking/latest/en/avatar_models/musetalk/)
⚠️ **归属说明**：这是 **OpenTalking 用 OmniRT 后端跑 MuseTalk** 的测量，不是原生 MuseTalk 脚本的测量；TTFV ≈ **1.77s（3090）/ 2.10s（4090）**，首轮总耗时 3.2–3.6s。

**⑤ [第三方实测] 4090 上「去掉 VAE 回传可到 60+fps，实际只有 30fps 左右」—— 定位到具体瓶颈行**
[MuseTalk issue #33「关于实时性的一些讨论」](https://github.com/TMElyralab/MuseTalk/issues/33) 用户 swx3027925806（2024-04-18）逐字：
> 「你好，在我们的的部署中发现一个问题，即在VAE的模型中，将GPU上的数据拷贝到CPU上花费了巨量的时间。简单来说就是**在不考虑这一步的情况下，实时性可以达到60+的fps。但是因为它的存在导致我们的性能只能在30fps左右**。请问有没有什么办法在这个基础上做到优化呢？这是因为显卡位宽所导致的吗？我们的实验环境是4090。」

其定位到的代码行为 `musetalk/models/vae.py` 的 `image = image.detach().cpu().permute(0, 2, 3, 1).float().numpy()`（逐字）。后续用户 hihowie 给出解释逐字：「这个时间不是拷贝数据到cpu的时间，是因为**gpu运算还未结束，会一直阻塞**，所以看起来拷贝时间很久」。
→ **这是解释官方「30fps+」与社区低帧率落差的最有技术价值的一条线索：瓶颈在 VAE decode + GPU→CPU 同步回传，而非 UNet 本身。**

**⑥ [第三方] 社区对「实时」的普遍质疑**：
- [issue #396「Has anyone actually been able to achieve real-time video generation?」](https://github.com/TMElyralab/MuseTalk/issues/396)（open）
- [issue #248「4090只能跑到5fps是怎么回事」](https://github.com/TMElyralab/MuseTalk/issues/248)（closed）
- [issue #49「为什么realtime脚本生成用时，比日志显示的fps实际要慢很多？」](https://github.com/TMElyralab/MuseTalk/issues/49)
- [issue #333「推理速度问题」](https://github.com/TMElyralab/MuseTalk/issues/333)、[issue #377「To achieve higher fps」](https://github.com/TMElyralab/MuseTalk/issues/377)、[issue #359「比MuseTalk1.0版本更慢了」](https://github.com/TMElyralab/MuseTalk/issues/359)
- [issue #173「Making MuseTalk 40% faster」](https://github.com/TMElyralab/MuseTalk/issues/173)
→ **结论：官方「30fps+」与大量社区实测（5–45fps）存在明显落差，选型时不要以 30fps+ 为基线。**

### (e) 流式 / 可打断

- **[官方] 有 realtime 推理脚本，但 60 秒级批处理语义为主**。README 提供 `sh inference.sh v1.5 realtime` 与 `python -m scripts.realtime_inference ...`；注意事项逐字：「1. Set `preparation` to `True` when processing a new avatar; 2. After preparation, the avatar will generate videos using audio clips from `audio_clips`; 4. Set `preparation` to `False` for generating more videos with the same avatar」。
- **[第三方] 所谓「实时」的口径问题**：issue #310 中楼主逐字反馈「想问一下实时推理的好像只有图片是实时生成的，音频是最后生成出来的」；维护者 zzzweakman 回复逐字：「是的，因为代码里有合成视频这一步，会将声音和图像序列合成视频」→ **音视频合成是事后步骤，不是逐帧流式**。
- **🔴 音频流式输入：官方明确表示未支持**。issue #33 用户 jinqinn（2024-05-06）逐字：
  > 「realtime-inference支持了图片的实时推理，但**不支持音频流失输入**，因musetalk引用了whisper组件，该组件好像不支持音频流失推理，请教下要如何进行调优呢？」

  同 issue 后续（2024-06-17）**项目组成员 czk32611 明确回复**逐字：
  > 「不好意思，**音频的流式处理我们暂时没有研究**。」

  → 这是**项目成员本人**对「MuseTalk 不支持音频流式输入」的确认。**音频必须先完整给出（或切块重跑），无法做到真正的边收边说**。对「实时交互数字人」场景而言，这是 hard blocker，也是为什么 LiveTalking 必须自行改造才能把它流式化。
  来源：[MuseTalk issue #33](https://github.com/TMElyralab/MuseTalk/issues/33)
- **可打断（barge-in）：未查到**。MuseTalk 仓库自身**没有**打断/打断重说机制；**该能力由 LiveTalking 封装提供**（见 LiveTalking Features 第 3 条）。
- 官方另有 `--skip_save_images` 加速选项（不落盘中间帧）。

### (f) 中文支持 / 半身-全身

- **中文：官方明确支持多语言**。README Overview 第 2 条逐字：「**supports audio in various languages, such as Chinese, English, and Japanese.**」
- **半身 / 全身：模型本身只处理面部区域，不决定全身**。Overview 第 1 条逐字：「modifies an unseen face according to the input audio, with a size of **face region of `256 x 256`**」；第 4 条逐字：「supports modification of the **center point of the face region** proposes, which **SIGNIFICANTLY** affects generation results」。
  → **MuseTalk 是「口型区域贴回原视频」的方案**：半身/全身取决于**你给的输入视频**，模型只改 256×256 面部区。**全身需靠上游框架（如 LiveTalking 的「全身视频拼接」）实现**，README 未提及全身。
- 音频编码使用冻结的 `whisper-tiny`，图像由冻结 VAE 编码，先验来自 `stable-diffusion-v1-4` 的 UNet 架构。

### (g) 已知坑：sm_120 / CUDA / torch / flash-attn / onnx

**🔴 坑 1（最关键）：Blackwell sm_120 需要 cu128，官方推荐的 cu118 直接报错**
issue #409 正文逐字给出对照表：

| PyTorch | CUDA | sm_120? |
|---|---|---|
| 2.6.0+cu124 | 12.4 | ❌ **no kernel image** errors |
| 2.10.0+cu126 | 12.6 | ❌ Same errors |
| **2.10.0+cu128** | **12.8** | ✅ **Works!** |

正文逐字：「MuseTalk recommends PyTorch 2.0.1+cu118, but Blackwell GPUs need **cu128+ for native sm_120 kernel support**.」
修复命令逐字：`pip install torch==2.10.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128`

**🔴 坑 2：mmpose/mmcv 在 Python 3.12 上装不上（无预编译 wheel + 源码编译失败）**
issue #409 逐字：「**mmcv has no pre-built wheels for Python 3.12 on any CUDA index. Building from source fails due to `pkg_resources` removal in Python 3.12.**」
作者的绕行方案（逐字，替换 `mmpose` 为 `mediapipe` + `face_alignment`）：`pip install mediapipe face-alignment`；用 mediapipe 的 478 点 face mesh 替代 mmpose 的 wholebody 模型；并映射 landmark 到 MuseTalk 的鼻梁索引：
> 「face_lm[28] = pts_478[6]（nose bridge top）；face_lm[29] = pts_478[197]（nose bridge mid）；face_lm[30] = pts_478[195]（nose bridge lower）」

**🔴 坑 3：上述 mediapipe→dlib 索引映射被后续实测质疑；mmcv 新版仍有 C++ 编译不兼容**
2026-08-05 用户 HeimdallCore 在同 issue 评论逐字：
> 「mmcv (we tried both **2.0.1** and the latest **2.2.0**) fails to compile against current PyTorch with an **ambiguous-overload C++ error around `Float8_e5m2fnuz`** operators — a genuine upstream incompatibility, not fixable with a flag.」
> 「a couple of the index mappings floating around, e.g. for points 28-30, **don't line up when we cross-checked them against independent mediapipe→dlib conversion tables**」
其更干净的方案逐字：「we replaced mmpose/dwpose landmark detection entirely with the standalone **face-alignment** PyPI package (MIT, FAN-based) — no MuseTalk dependency, no CUDA/C++ build step, and it returns native **68-point dlib-format** landmarks directly」；补丁仅 2 处（`musetalk/utils/preprocessing.py`），并已端到端验证 361 帧。
→ **对 5060 Ti + Python 3.12 的部署建议：直接采用 face-alignment 方案，避开 mmcv 编译。**

**🔴 坑 4：50 系显卡支持的社区实操记录（issue #334）**
- 楼主 wilsonlv 逐字：「现在不支持50系列显卡，希望可以兼容一下」
- wanlichina 逐字：「可以兼容的，自动动手装最新的 **pytorch 2.7+cu128**，修改模型的 `torch.load` 的加载参数，你跟着新环境的运行错误去找位置就好」
- 随后 `mim install mmcv==2.0.1` 在 cu128 上编译失败（705 行报错），wanlichina 逐字修正：「**MMCV 2.0.1不支持cu128，你改成 mim install mmvc==2.1.0 就好了**」；但升级后出现 `ImportError: DLL load failed while importing _ext`（Windows）。
- codestart-zhu 逐字确认：「这个我也遇到过，改了mmcv的版本，环境是 **pytorch2.7+cu128，现在5090能跑起来的**」
来源：[MuseTalk issue #334「50系列显卡支持」](https://github.com/TMElyralab/MuseTalk/issues/334)
→ **要点：`torch.load` 加载参数需按新 PyTorch 调整（新版默认 `weights_only=True` 会拒绝加载旧 checkpoint）。**

**坑 5：推理未加 no_grad 导致显存暴涨（与 LiveTalking 同源）**
[MuseTalk PR #349](https://github.com/TMElyralab/MuseTalk/pull/349) 标题逐字：「Fix issue #235: Use `torch.no_grad()` in inference to prevent excessive…」（已 merged）。LiveTalking 侧对应的就是 PR #612（17.62GB→3.9GB）。

**坑 6：依赖版本锁定陈旧**
`requirements.txt` 逐字：`diffusers==0.30.2`、`accelerate==0.28.0`、`numpy==1.23.5`、`tensorflow==2.12.0`、`opencv-python==4.9.0.80`、`transformers==4.39.2`、`huggingface_hub==0.30.2`、`librosa==0.11.0`、`einops==0.8.1`、`gradio==5.24.0`。
- **含 `tensorflow==2.12.0`**：运行时会打印 `TF-TRT Warning: Could not find TensorRT`（见 issue #310 日志），属噪音非错误。
- **无 `flash-attn`、无 `triton`、无 `onnx`/`tensorrt` 依赖** → **不存在 flash-attn/triton 编译坑，也无 ONNX/TensorRT 版本坑**（上表 OpenTalking 的 `OmniRT` 是第三方自建后端，不是 MuseTalk 官方能力）。
- **numpy 1.23.5 与 torch 2.10 组合需实测**，issue #409 的成功组合未列 numpy 版本。
- 上游 HF 权重最后更新时间为 **2025-03-31**（`lastModified`），仓库代码在 2026 年主要由社区 issue/PR 驱动维护。

---

## 交叉结论速览（C1 三仓库）

| 项目 | 推理峰值显存（最可信数字） | 16GB 单卡推理 | 训练显存 | 实时 FPS（实测/官方） | 首帧延迟 | 流式/打断 | 中文 | sm_120 状态 |
|---|---|---|---|---|---|---|---|---|
| **LiveTalking** | musetalk **3.9GB**（修 no_grad 后，PR#612 实测 5090）；未修时 17.62GB→OOM；wav2lip ~200MB（作者估算） | ✅ 充足（务必含 PR#612） | — （不训练，用现成权重） | wav2lip256 3060 **60fps**；musetalk 3080Ti **42** / 3090 **45** / 4090 **72** [官方] | 未查到（WebRTC 建链 6s~数分钟，非模型延迟） | ✅ **支持打断**、WebRTC/RTMP/多并发 | ✅ 完整 | 无 sm_120 issue（total 0）；官方验证 torch 2.9.1+cu128 |
| **Ultralight-Digital-Human** | **未查到实测值**（模型移动端级，LiveTalking 侧估 ~1.5GB FP32 [他人估算]） | ✅ 推断充足 | 未查到（默认 batch=1） | **<10ms/帧**（2080 多并发，需转 ONNX）[作者实测] | 未查到 | ⚠️ 支持但**代码未完善**；继任 FeatherTalk 已开源 C++ 流式 | ✅ 完整（康辉 demo） | 🔴 官方栈 torch1.13.1+cu117 **无 sm_120 内核**，必须自行迁移 cu128 |
| **MuseTalk** | **~3.6GB**（RTX 5060 Ti 16GB 实测，issue #409）；OpenTalking 峰值 **5.078–5.203GB**；社区实测可压到 **11GB** 以下 | ✅ 充足；**训练不行**（stage1 ≥32GB / stage2 ≥54GB，需 8×H20） | 官方：stage1 bs8 ≈32GB、bs32 ≈74GB；stage2 ≈54–85GB | 官方 **30fps+ on V100**；5060Ti 实测 **≈5.8fps** [推算]；OpenTalking 3090 **28.87** / 4090 **24.77** | TTFV **1.77s**（3090）/ **2.10s**（4090）[OpenTalking 第三方]；首轮总 3.24–3.61s | ❌ 自身无打断；**项目成员确认「音频的流式处理我们暂时没有研究」**；打断靠 LiveTalking | ✅ 官方声明中/英/日 | 🔴 torch 需 **cu128**（cu124/cu126 均报 no kernel image）；🔴 **mmcv/mmpose 在 Py3.12 装不上/编译失败**，建议改用 face-alignment |

### 针对 RTX 5060 Ti 16GB / sm_120 的三条硬结论
1. **MuseTalk 有同款硬件实测**（issue #409，2026-03-18 + 2026-08-05 追加）：16GB 下 **~3.6GB VRAM** 跑通 361 帧、口型正确；**但吞吐约 5.8fps，达不到实时**，且必须用 **torch 2.10.0+cu128** 且绕开 mmcv（用 face-alignment）。
2. **LiveTalking 是目前唯一「原生实时 + 可打断 + 中文 + 全身」的整合方案**，官方验证栈正是 **CUDA 12.8**（与 sm_120 对齐），且 sm_120 无已知 issue；**但 musetalk 后端必须包含 2026-08-20 合并的 no_grad 修复（PR #612），否则 16GB 卡按默认 batch=16 会 OOM**。想稳，用 `--model wav2lip`（3060 都能 60fps）或把 batch 降到 8。
3. **Ultralight-Digital-Human 是三仓库中 16GB 适配风险最高的**：官方栈（torch 1.13.1 / cu117 / py3.10）**在 Blackwell 上不可用**，且流式代码长期未开源、推理默认跑 CPU。作者继任项目 **FeatherTalk**（C++ 流式、FeatherHuBERT 音频侧算力降 40×、audio+UNet <1M）更值得跟进。

### 补充：维护活跃度是三仓库最被低估的选型维度
截至 **2026-09-22**，最后提交时间分别为：LiveTalking **2026-09-13**（9 天前）、Ultralight-Digital-Human **2026-07-22**、MuseTalk **2025-09-26**（近一年前）。
sm_120 是 2025 年后的新平台，**其适配问题只能靠上游持续修**。MuseTalk 上游已停滞，2026 年的 Blackwell 适配案例全部来自社区自发贡献（#409、#334）且**尚未合并进主干**——选择 MuseTalk 必须自行维护补丁分支。LiveTalking 在 2026 年 8–9 月仍有合并提交，是其被推荐为首选整合层的额外理由。

---

## 来源清单

**官方仓库 / 文档**
1. https://github.com/lipku/LiveTalking
2. https://raw.githubusercontent.com/lipku/LiveTalking/main/README.md
3. https://raw.githubusercontent.com/lipku/LiveTalking/main/requirements.txt
4. https://doc.livetalking.ai/docs/faq/
5. https://github.com/anliyuan/Ultralight-Digital-Human
6. https://raw.githubusercontent.com/anliyuan/Ultralight-Digital-Human/master/README.md （main 分支 README 为 404）
7. https://github.com/anliyuan/FeatherTalk （继任项目，README 内链）
8. https://github.com/TMElyralab/MuseTalk
9. https://raw.githubusercontent.com/TMElyralab/MuseTalk/main/README.md
10. https://raw.githubusercontent.com/TMElyralab/MuseTalk/main/requirements.txt
11. https://huggingface.co/TMElyralab/MuseTalk
12. https://arxiv.org/abs/2410.10122 （MuseTalk 技术报告）

**GitHub Issues / PR（含实测数字）**
13. https://github.com/lipku/LiveTalking/pull/612 （musetalk 实时推理显存 17.6G→3.9G，2026-08-20 merged）
14. https://github.com/lipku/LiveTalking/issues/192 （4090 musetalk OOM，17.77GiB allocated）
15. https://github.com/lipku/LiveTalking/issues/368 （WebRTC 建链 6s~数分钟）
16. https://github.com/lipku/LiveTalking/issues/618 （中文 ASR/TTS 后端请求，2026 活跃）
17. https://github.com/lipku/LiveTalking/issues/535 （多并发会话）
18. https://github.com/TMElyralab/MuseTalk/issues/409 （**RTX 5060 Ti 16GB / sm_120 实测：~3.6GB VRAM，7s 音频→30s 推理**；2026-03-18 及 2026-08-05 追加）
19. https://github.com/TMElyralab/MuseTalk/issues/392 （VRAM 需求；GTX 1060 3GB 可跑但极慢）
20. https://github.com/TMElyralab/MuseTalk/issues/310 （4090D 实时爆显存；实测可压到 11G；batch=15 约 20G）
21. https://github.com/TMElyralab/MuseTalk/issues/334 （50 系支持；mmcv 2.0.1 不支持 cu128 → 2.1.0；5090 可跑）
22. https://github.com/TMElyralab/MuseTalk/issues/362 （CUDA out of memory）
23. https://github.com/TMElyralab/MuseTalk/issues/331 （OutOfMemoryError）
24. https://github.com/TMElyralab/MuseTalk/issues/396 （社区质疑实时性）
25. https://github.com/TMElyralab/MuseTalk/issues/248 （4090 只能 5fps）
26. https://github.com/TMElyralab/MuseTalk/issues/49 （realtime 脚本实际比日志 fps 慢）
27. https://github.com/TMElyralab/MuseTalk/issues/333 （推理速度问题）
28. https://github.com/TMElyralab/MuseTalk/issues/377 （To achieve higher fps）
29. https://github.com/TMElyralab/MuseTalk/issues/359 （比 1.0 更慢）
30. https://github.com/TMElyralab/MuseTalk/issues/173 （Making MuseTalk 40% faster）
31. https://github.com/TMElyralab/MuseTalk/pull/349 （推理加 torch.no_grad() 修复）
32. https://github.com/anliyuan/Ultralight-Digital-Human/issues/120 （推理内存占用近 10G）33. https://github.com/anliyuan/Ultralight-Digital-Human/issues/74 （推理全程走 CPU，GPU 无动静）
34. https://github.com/anliyuan/Ultralight-Digital-Human/issues/103 （催开源实时流代码）
35. https://github.com/anliyuan/Ultralight-Digital-Human/issues/101 （音频特征提取太慢）
36. https://github.com/anliyuan/Ultralight-Digital-Human/issues/161 （Whisper-tiny vs WeNet 速度）
37. https://github.com/anliyuan/Ultralight-Digital-Human/issues/18 （200 epochs 效果抖）
38. https://github.com/anliyuan/Ultralight-Digital-Human/issues/81 （移动端 wenet 特征提取位置）
39. https://github.com/anliyuan/Ultralight-Digital-Human/issues/226 （FunASR/SenseVoice 集成建议）
40. https://github.com/anliyuan/Ultralight-Digital-Human/issues/52 （hubert 报错）

**第三方博客 / 文档 / 基准**
41. https://www.cnblogs.com/livetalking/articles/22134438 （拆解 LiveTalking 资源消耗；三模型显存估算 + FPS 表 + 40ms 预算；作者=lipku）
42. https://datascale-ai.github.io/opentalking/latest/en/avatar_models/musetalk/ （MuseTalk 峰值显存 5.078/5.203GB、TTFV 1769/2095ms、Steady FPS 28.87/24.77）
43. https://datascale-ai.github.io/opentalking/latest/en/benchmark/ （OpenTalking 总基准页）
44. https://deepwiki.com/anliyuan/Ultralight-Digital-Human/3.2-digital-human-model-training （Ultralight 训练参数：epochs 200 / batchsize 1；SyncNet 参数已过时）
45. https://hf-mirror.com/TMElyralab/MuseTalk/resolve/main/musetalkV15/unet.pth （HEAD 实测 Content-Length 3,400,074,924 B）
46. https://hf-mirror.com/TMElyralab/MuseTalk/resolve/main/musetalk/pytorch_model.bin （HEAD 实测 Content-Length 3,400,076,549 B）

**2026 年状态核实（最后提交 / Star / License）—— 无 API rate limit 的取数端点**
47. https://github.com/lipku/LiveTalking/commits/main.atom （`<updated>` 首条 = **2026-09-13T11:56:50Z**）
48. https://github.com/anliyuan/Ultralight-Digital-Human/commits/master.atom （`<updated>` 首条 = **2026-07-22T12:29:23Z**；同时验证 `main` 分支不存在）
49. https://github.com/TMElyralab/MuseTalk/commits/main.atom （`<updated>` 首条 = **2025-09-26T05:44:17Z**）
50. https://img.shields.io/github/stars/lipku/LiveTalking.json （9.6k）· https://img.shields.io/github/license/lipku/LiveTalking.json （Apache-2.0）
51. https://img.shields.io/github/stars/anliyuan/Ultralight-Digital-Human.json （2.6k）· https://img.shields.io/github/license/anliyuan/Ultralight-Digital-Human.json （not specified）
52. https://img.shields.io/github/stars/TMElyralab/MuseTalk.json （6.6k）· https://img.shields.io/github/license/TMElyralab/MuseTalk.json （not identifiable by github）
53. https://img.shields.io/github/stars/anliyuan/FeatherTalk.json （93）· https://img.shields.io/github/last-commit/anliyuan/FeatherTalk.json （2026-09）

**新增 issue 证据**
54. https://github.com/TMElyralab/MuseTalk/issues/33 （**4090：去掉 VAE→CPU 回传可达 60+fps，实际仅约 30fps**；项目成员 czk32611 确认「**音频的流式处理我们暂时没有研究**」）
55. https://www.cnblogs.com/livetalking/articles/22134438 （补充：avatar 帧缓存 **10s 1080p ≈ 1.61GB RAM**；默认 max_session **5 路**；**NVENC 并发上限表**；单路 1080p libx264 软编 ≈ **0.45 核**；反压机制 `buffer_size >= 5`）

**已尝试但未能取用（记录以免重复劳动）**
56. https://blog.csdn.net/lipku/article/details/148593974 （作者 lipku 的「实时数字人多并发」，CSDN 返回 2150 字节反爬壳页，正文未取到）
57. https://blog.csdn.net/beautifulmemory/article/details/152012956 （LiveTalking Linux 搭建教程，同样被 CSDN 反爬拦截，2274 字节）
58. https://blog.gitcode.com/3fe07b50fbd0927c431432a2e3d7d359.html （MuseTalk 4090 性能优化，页面自述由 **AIGC 生成**且「登录后查看全文」，不具引用价值，已弃用）
59. https://raw.githubusercontent.com/anliyuan/FeatherTalk/main/README.md 与 `/master/README.md` （均返回 0 字节；FeatherTalk 的显存/延迟数字**未查到**，其能力描述仅能从 Ultralight README 转述）

60. LICENSE 核验：`https://raw.githubusercontent.com/lipku/LiveTalking/main/LICENSE` → **HTTP 200**（Apache License 正文实取）；`.../anliyuan/Ultralight-Digital-Human/master/LICENSE` 与 `.../TMElyralab/MuseTalk/main/LICENSE` → **HTTP 000**（连接失败，非 404），**故二者的许可证状态未能确认**，仅以 shields.io 结果为准。

**检索方式说明**：GitHub issue/PR 正文通过 github.com 页面内嵌 JSON（`bodyHTML`）解析获取。**本次调研中途 `api.github.com` 未认证额度（core 60/hr，与同 IP 的其他 agent 共享）被耗尽**，故后续改用：`commits/<branch>.atom`（提交时间）、shields.io（Star/License）、`raw.githubusercontent.com`（README）、以及 `web_search` + curl 目标页。所有数字均逐字取自上述来源，未做任何推测性填补；无来源处一律标注「未查到」。
