findings_C1.md

实时数字人开源项目 · 16GB 显卡选型调研(C1 分片)

目标硬件:RTX 5060 Ti 16GB(sm_120 / Blackwell)、32GB RAM、Ubuntu、NVIDIA 驱动 610.43.02、CUDA 13.2 调研日期:2026-09-22 责任人:C1 范围:仅 3 个仓库(lipku/LiveTalking、anliyuan/Ultralight-Digital-Human、TMElyralab/MuseTalk)

标注约定:[作者自述] = 项目作者/官方文档;[第三方实测] = 用户在 issue/博客的实测;[推算] = 由可验证数字推导,非直接引用;未查到 = 未找到可验证来源。


0. 2026 年状态核实(最后提交时间 / Star / License)

测量方式:curl -sL "https://github.com/<owner>/<repo>/commits/<branch>.atom" 取 <updated> 首条(提交时间戳,无 API rate limit);Star/License 取 shields.io img.shields.io/github/{stars,license,last-commit}/<owner>/<repo>.json。核实日期:2026-09-22。

仓库默认分支最后提交时间StarLicense
lipku/LiveTalkingmain(master 不存在)2026-09-13T11:56:50Z(距调研日仅 9 天)9.6kApache-2.0
anliyuan/Ultralight-Digital-Humanmaster(main 不存在)2026-07-22T12:29:23Z2.6kshields.io 返回 not specified(README 徽章指向 ./LICENSE,但本次 curl 该文件返回 000/连接失败而非 404,许可证状态存疑,未验证;对比 LiveTalking 的 LICENSE 返回 200 + Apache License 正文)
TMElyralab/MuseTalkmain(master 不存在)2025-09-26T05:44:17Z(⚠️ 已近一年未更新)6.6kshields.io 返回 not identifiable by github(HF 页标注 creativeml-openrail-m)
anliyuan/FeatherTalk(Ultralight 继任者,参考项)—2026-09(shields "september")93Apache-2.0

结论 / 对选型的影响:

  1. LiveTalking 是三者中唯一在 2026 年持续高频维护的项目(最后提交 2026-09-13,且 PR #612 于 2026-08-20 合并)。目标硬件 sm_120 属于新平台,维护活跃度是能否踩到坑并被修好的关键变量,此项 LiveTalking 明显占优。
  2. MuseTalk 上游代码已停滞近一年(2025-09-26 之后无提交)。其 2026 年的 sm_120 适配进展全部来自社区 issue/PR 自述(如 #409、#334),不存在官方修复承诺。选它需接受"自己打补丁"的现实。
  3. Ultralight-Digital-Human 在 2026 年仍有更新(2026-07-22),但作者已把重心转向 FeatherTalk;本仓库的「流式代码未完善」缺口在其继任者中才补齐。
  4. 默认分支陷阱:Ultralight 默认分支是 master,其 main 分支不存在(README.md 请求返回 404,仅 14 字节)。自动化脚本若硬编码 main 会拿到 404 —— 本报告初次抓取即撞上此坑。

1. lipku/LiveTalking

定位:实时交互流式数字人引擎(不是单一模型,而是 wav2lip / musetalk / Ultralight 三模型的流式封装 + WebRTC/RTMP 推流管线)。

(c) 显存需求 + 16GB 单卡能否跑

[作者自述] 官方性能表(README「6. 性能指标」):需要区分「模型本体显存」与「推理峰值显存」。作者本人(博客园 ID「恒中」= lipku)在博客中给出了按源码估算的三模型显存:

模型技术路线显存占用(估算)精度输入分辨率
Wav2LipCNN 编解码 + 空间注意力~200MBFP32256x256
MuseTalkUNet (Diffusion) + VAE + Whisper~1.2GB(FP32)/ ~600MB(FP16)FP16256×256
UltraLight轻量 UNet + HuBERT~1.5GBFP32160×160

原文逐字引用:「MuseTalk | UNet (Diffusion) + VAE + Whisper | ~1.2GB(FP32)/ ~600MB(FP16) | FP16 | 256×256」 来源:拆解 LiveTalking 的资源消耗:GPU、CPU、内存、带宽到底花在哪? ⚠️ 注意:该表标注为「显存占用(估算)」,是作者读源码的估算值,不是实测峰值。

[第三方实测] musetalk 真实峰值显存 —— 17.62 GB → 3.9 GB(这条最关键): PR #612 标题逐字:「fix: MuseReal.inference_batch 缺 @torch.no_grad(),实时推理显存 17.6G→3.9G」,已于 2026-08-20 被作者 lipku 合并(MERGED)。PR 正文逐字引用:

「问题:同文件的 warm_up / init 都带 no_grad,唯独实时推理路的 inference_batch 漏了,调用方 BaseAvatar.inference 也没有 ⇒ 每批推理都在构建 autograd 图,显存被求导图白白占掉。 实测(RTX 5090,musetalk,batch_size=16):torch allocated 峰值 17.62 GB → …」

来源:lipku/LiveTalking PR #612(第四方摘录自 PR 页 twitter:description 元数据,正文尾部被页面截断,但两个数字完整可读)

[第三方实测] 未修复时 4090 直接 OOM(issue #192,2024-08-07):

「RuntimeError: CUDA out of memory. Tried to allocate 256.00 MiB (GPU 0; 23.65 GiB total capacity; 17.77 GiB already allocated; 45.19 MiB free; 18.47 GiB reserved in total by PyTorch)」

社区给出的解法(issue #192 评论,leos-code)逐字:

「在 musereal.py 的 inference方法上添加 @torch.no_grad() 注解;启动app.py时,设置batch_size为8或更小,我设置8可以正常运行」

来源:LiveTalking issue #192「musetalk 启动后 oom」

[作者自述] 内存(RAM)与并发上限 —— 与显存同等重要的是内存(来源同上博客园文章): 作者从源码逐字拆解:系统启动时会把整个 avatar 视频所有帧加载进内存。以 10 秒 1080p 视频(25fps) 为例:

「总帧数:10s × 25fps = 250 帧 全帧缓存:250 × 1920 × 1080 × 3 字节 = 250 × 6.2 MB ≈ 1.56 GB 人脸缓存:250 × 256 × 256 × 3 字节 = 250 × 0.2 MB ≈ 49 MB 坐标数据:250 × 4 × 4 字节 ≈ 4 KB(可忽略) ───────────────────────────────────────────── 总计:约 1.61 GB(仅一个avatar视频)」

并发默认值逐字:「5 路(默认 max_session)」。→ 32GB RAM 的机器上,多路数字人时 avatar 帧缓存是主要内存开销(每路约 1.6GB/10 秒 1080p 素材),显存反而不是瓶颈。

NVENC 硬编并发上限(若把 libx264 软编换成 h264_nvenc)(同来源):

GPUNVENC 芯片数并发编码上限说明
RTX 30601 个3 路消费卡驱动限制
RTX 3080Ti / 30901 个3 路消费卡驱动限制
RTX 40902 个5-8 路双编码器 + 新驱动放宽
RTX A5000 / A60001-2 个20+ 路专业卡无驱动限制
A100 / H100无 NVENC0 路纯计算卡,没有编码器!

⚠️ 关键提醒(逐字):「A100/H100 是纯 CUDA 计算卡,没有 NVENC 编码器。如果选用 A100 做高并发推理,H.264 编码依然要靠 CPU 软编。」 → RTX 5060 Ti 未在该表中,按消费卡同类推断为 1 个 NVENC / 3 路并发上限([推算],表中无 50 系数据,属未验证项)。默认走 libx264 纯软件编码,作者测算单路 1080p 推流持续占用约 0.45 个 x86 CPU 核心,5 路约 2.0-2.5 核。

[作者自述] 反压机制:输出缓冲区积压超过 5 帧时主动休眠(base_avatar.py:482-484):if buffer_size >= 5: time.sleep(0.04 * buffer_size * 0.8)。每会话 4 线程(render() / inference() / process_frames() / process_tts())。

16GB 单卡结论:

(d) 端到端延迟:首帧延迟 + 推理 FPS

[作者自述] 官方 FPS 表(README「6. 性能指标」,逐字):

模型显卡FPS
wav2lip256RTX 306060
wav2lip256RTX 3080Ti120
musetalkRTX 3080Ti42
musetalkRTX 309045
musetalkRTX 409072

官方补充逐字:「后端日志 inferfps = GPU 推理帧率, finalfps = 最终推流帧率,两者均需 >=25 才算实时」。 推荐配置逐字:「wav2lip256 推荐 RTX 3060 及以上;musetalk 推荐 RTX 3080Ti 及以上」。

[作者自述] 每帧 40ms 预算:「每一帧只有 40ms 的预算(25fps)。这个流水线上的任何一个环节超时,用户就会感知到卡顿。」(同上博客园文章)

首帧延迟(first-frame latency):未查到官方数字。 最接近的第三方现象记录是 issue #368,用户描述 WebRTC 首画建立耗时波动极大:

「webrtc从点击连接到前端有画面显示的过程中,有时候超级慢要几分钟,有时候只要6s的样子」

来源:LiveTalking issue #368(注意:6s 是 WebRTC 链路建立耗时,不是模型首帧推理耗时,不可当模型延迟引用)

⚠️ 官方 FPS 表未标注测试分辨率,也未覆盖 sm_120 显卡(无 50 系数据)。

(e) 流式 / 可打断

[作者自述] 支持,且是核心卖点(README「Features」逐字):

「支持打断重说」在「AI 数字人客服」场景中再次明确。架构上 /human(文本,echo/chat 模式)、/humanaudio(音频文件直接播放)接口,每个连接分配唯一 sessionid 支持多用户并发。

(f) 中文支持 / 半身-全身

(g) 已知坑:sm_120 / CUDA / torch / flash-attn / onnx


2. anliyuan/Ultralight-Digital-Human

定位:超轻量级数字人,主打移动端实时。README 明确其为「personal」路线(每人一段视频训一个专属模型)。作者已开新项目 FeatherTalk(C++ 流式推理)作为继任者。 ⚠️ 默认分支为 master(main 分支下 README 返回 404)。

(c) 显存需求 + 16GB 单卡能否跑

显存(VRAM)实测数字:未查到。仓库 README、作者文档均未给出任何显存数字;也未找到第三方实测报告。

可用的周边证据:

16GB 单卡结论:显存侧必然够用(模型是移动端级别),但没有可引用的实测显存数字,报告中应标注「未查到实测值,推断极小」。真正的限制在软件栈兼容性(见 (g))与内存占用。

(d) 端到端延迟:首帧延迟 + 推理 FPS

⚠️ 重大坑(会直接毁掉实时性体验)(issue #74 逐字):

「推理时,使用文档中的语句:python inference.py ... 但是全部用的是CPU去处理,GPU没动静,是需要配置参数吗」

来源:issue #74「推理时能用GPU吗」 → 官方 inference.py 默认跑在 CPU 上,需自行改 device。

(e) 流式 / 可打断

(f) 中文支持 / 半身-全身

(g) 已知坑:sm_120 / CUDA / torch / flash-attn / onnx


3. TMElyralab/MuseTalk

定位:音频驱动口型同步(lip-sync)模型,在 ft-mse-vae latent 空间做单步 inpainting,不是 diffusion 模型。被 LiveTalking 作为三种可选后端之一封装。 ⭐ 本项目有 RTX 5060 Ti 16GB / sm_120 的直接实测报告(issue #409),是本报告最有价值的证据。

(c) 显存需求 + 16GB 单卡能否跑

① [第四方实测 · 目标硬件直接命中] RTX 5060 Ti 16GB 上 ~3.6GB VRAM MuseTalk issue #409 标题逐字:「Field Report: MuseTalk V1.5 working on RTX 5060 Ti (Blackwell sm_120) with Python 3.12 + mediapipe patch」(2026-03-18,作者 chefboyrdave21)。正文「Performance」段落逐字:

「- 7sec audio → 30sec inference → MP4 output

  • ~3.6GB VRAM for MuseTalk models
  • Reference image face detection cached after first call」

文末环境逐字:「Setup: RTX 5060 Ti 16GB | Ubuntu 24.04 | Python 3.12.3 | PyTorch 2.10.0+cu128 | MuseTalk V1.5」

来源:TMElyralab/MuseTalk issue #409

② [同一 issue 内 2026-08-05 追加实测] 用户 HeimdallCore 在同 issue 评论(2026-08-05)逐字:

「We ran into the same Blackwell (RTX 5060 Ti, sm_120) situation... We verified this end-to-end (real MuseTalk V1.5 inference, 361 frames, correct lip-sync in the output).」

→ 361 帧端到端跑通,进一步确认 16GB sm_120 可跑。

③ [官方自述] 最低 4GB VRAM(fp16) README「Gradio Demo」段落逐字:

「For minimum hardware requirements, we tested the system on a Windows environment using an NVIDIA GeForce RTX 3050 Ti Laptop GPU with 4GB VRAM. In fp16 mode, generating an 8-second video takes approximately 5 minutes.」

来源:MuseTalk README(下载页:README 原文)

④ [第三方实测] 3GB 老卡也能跑(极慢) issue #392 逐字:

「I have only a gtx 1060 3gb PC. Is that enough to run this model?」 「I was able to run musetalk with a gtx 1060 3gb. It's just ultra slow.」

来源:MuseTalk issue #392「What are the VRAM requirements」

⑤ [第三方实测] 实时推理 11GB 足够,瓶颈是算力而非显存 issue #310 用户 wanlichina 逐字:

「4090D的运算能力,你哪怕改成4或者2都能满足你的试试推理要求。显存占用我测试下来,能做到11G。也就是3080,4080,5080都能跑,但是瓶颈不在显存,在GPU算力,因为GPU占用率爆了,多开满足不了实时推理性能」

同 issue 楼主 codestart-zhu 未调 batch 时逐字:torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 400.00 MiB (GPU 0; 23.64 GiB total capacity; 22.15 GiB already allocated...);调到 batch_size=15 后「现在占用在20g左右」。 来源:MuseTalk issue #310「实时推理4090d爆显存问题」

⑥ [推算] fp16 权重体积 —— 由 HuggingFace 上文件实际字节数推导

文件实测字节数换算fp16 推算
musetalkV15/unet.pth3,400,074,924 B3.40 GB / 3.17 GiB≈1.70 GB(fp16)
musetalk/pytorch_model.bin(v1.0)3,400,076,549 B3.40 GB / 3.17 GiB≈1.70 GB(fp16)

测量方式:对 https://hf-mirror.com/TMElyralab/MuseTalk/resolve/main/... 发 HTTP HEAD 读取 Content-Length。 → 3.4GB ≈ SD1.4 UNet 的 FP32 体积(~860M 参数 × 4B)。因此发布的 unet.pth 是 fp32 权重,fp16 约为其一半 ≈ 1.7GB。此为推算值,非官方标注。 (HF 官方仓库文件清单见 huggingface.co/TMElyralab/MuseTalk,文件仅为 musetalkV15/unet.pth、musetalkV15/musetalk.json、musetalk/pytorch_model.bin、musetalk/musetalk.json。)

⑦ [官方自述] 训练显存 —— 16GB 单卡完全不可能训练 README「GPU Memory Requirements」逐字(基于 8× NVIDIA H20 测试):

Stage 1Batch SizeGrad AccumMemory per GPU
81~32GB
161~45GB
321~74GB ✓
Stage 2Batch SizeGrad AccumMemory per GPU
18~54GB
22~80GB
28~85GB ✓

16GB 单卡结论:

(d) 端到端延迟:首帧延迟 + 推理 FPS

① [官方自述] 30fps+ on Tesla V100 —— README 出现两次,逐字:

② [第三方实测 · 目标硬件] RTX 5060 Ti 16GB:约 5.8 fps(非实时) 由 issue #409 逐字数据推算:7sec audio → 30sec inference。音频 7 秒 @25fps = 175 帧,30 秒推理 → 175/30 ≈ 5.8 fps([推算],基于第三方实测原值)。这是本报告中唯一的目标显卡直接可算的吞吐数据,且结论是「远未达到实时」。 ⚠️ 需注意该方法用的是批量推理而非 realtime_inference.py 流式脚本,且原文未说明是否启用 fp16/优化。

③ [第三方实测] LiveTalking 封装下的 musetalk FPS(见上文 LiveTalking 官方表):RTX 3080Ti 42 / RTX 3090 45 / RTX 4090 72。

④ [第三方 ] 首帧延迟(TTFV)实测 —— 这是本项目唯一可引用的首帧数字 OpenTalking 官方 benchmark 的 MuseTalk 页表格(该页说明逐字:「The numbers below are summarized from Benchmark. Steady FPS is model-generation throughput, not full user-perceived latency; STT, LLM, TTS, queueing, and WebRTC still affect the complete experience.」):

HardwareBackendOutputSteady FPSFirst-turn total/msTTFV/msPeak inference VRAM/GB
RTX 3090OmniRT512×512 / 25fps28.8683235.5181769.4845.078
RTX 4090OmniRT512×512 / 25fps24.7673605.5642095.5225.203
NPU 910B2OmniRT512×512 / 25fps12.2765781.4534211.7218.754

逐字引用:「Peak inference VRAM/GB | 5.078」「TTFV/ms | 1769.484」「First-turn total/ms | 3235.518」 来源:OpenTalking · MuseTalk 文档 ⚠️ 归属说明:这是 OpenTalking 用 OmniRT 后端跑 MuseTalk 的测量,不是原生 MuseTalk 脚本的测量;TTFV ≈ 1.77s(3090)/ 2.10s(4090),首轮总耗时 3.2–3.6s。

⑤ [第三方实测] 4090 上「去掉 VAE 回传可到 60+fps,实际只有 30fps 左右」—— 定位到具体瓶颈行 MuseTalk issue #33「关于实时性的一些讨论」 用户 swx3027925806(2024-04-18)逐字:

「你好,在我们的的部署中发现一个问题,即在VAE的模型中,将GPU上的数据拷贝到CPU上花费了巨量的时间。简单来说就是在不考虑这一步的情况下,实时性可以达到60+的fps。但是因为它的存在导致我们的性能只能在30fps左右。请问有没有什么办法在这个基础上做到优化呢?这是因为显卡位宽所导致的吗?我们的实验环境是4090。」

其定位到的代码行为 musetalk/models/vae.py 的 image = image.detach().cpu().permute(0, 2, 3, 1).float().numpy()(逐字)。后续用户 hihowie 给出解释逐字:「这个时间不是拷贝数据到cpu的时间,是因为gpu运算还未结束,会一直阻塞,所以看起来拷贝时间很久」。 → 这是解释官方「30fps+」与社区低帧率落差的最有技术价值的一条线索:瓶颈在 VAE decode + GPU→CPU 同步回传,而非 UNet 本身。

⑥ [第三方] 社区对「实时」的普遍质疑:

(e) 流式 / 可打断

(f) 中文支持 / 半身-全身

(g) 已知坑:sm_120 / CUDA / torch / flash-attn / onnx

🔴 坑 1(最关键):Blackwell sm_120 需要 cu128,官方推荐的 cu118 直接报错 issue #409 正文逐字给出对照表:

PyTorchCUDAsm_120?
2.6.0+cu12412.4❌ no kernel image errors
2.10.0+cu12612.6❌ Same errors
2.10.0+cu12812.8✅ Works!

正文逐字:「MuseTalk recommends PyTorch 2.0.1+cu118, but Blackwell GPUs need cu128+ for native sm_120 kernel support.」 修复命令逐字:pip install torch==2.10.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128

🔴 坑 2:mmpose/mmcv 在 Python 3.12 上装不上(无预编译 wheel + 源码编译失败) issue #409 逐字:「mmcv has no pre-built wheels for Python 3.12 on any CUDA index. Building from source fails due to pkg_resources removal in Python 3.12.」 作者的绕行方案(逐字,替换 mmpose 为 mediapipe + face_alignment):pip install mediapipe face-alignment;用 mediapipe 的 478 点 face mesh 替代 mmpose 的 wholebody 模型;并映射 landmark 到 MuseTalk 的鼻梁索引:

「face_lm[28] = pts_478[6](nose bridge top);face_lm[29] = pts_478[197](nose bridge mid);face_lm[30] = pts_478[195](nose bridge lower)」

🔴 坑 3:上述 mediapipe→dlib 索引映射被后续实测质疑;mmcv 新版仍有 C++ 编译不兼容 2026-08-05 用户 HeimdallCore 在同 issue 评论逐字:

「mmcv (we tried both 2.0.1 and the latest 2.2.0) fails to compile against current PyTorch with an ambiguous-overload C++ error around Float8_e5m2fnuz operators — a genuine upstream incompatibility, not fixable with a flag.」 「a couple of the index mappings floating around, e.g. for points 28-30, don't line up when we cross-checked them against independent mediapipe→dlib conversion tables」

其更干净的方案逐字:「we replaced mmpose/dwpose landmark detection entirely with the standalone face-alignment PyPI package (MIT, FAN-based) — no MuseTalk dependency, no CUDA/C++ build step, and it returns native 68-point dlib-format landmarks directly」;补丁仅 2 处(musetalk/utils/preprocessing.py),并已端到端验证 361 帧。 → 对 5060 Ti + Python 3.12 的部署建议:直接采用 face-alignment 方案,避开 mmcv 编译。

🔴 坑 4:50 系显卡支持的社区实操记录(issue #334)

坑 5:推理未加 no_grad 导致显存暴涨(与 LiveTalking 同源) MuseTalk PR #349 标题逐字:「Fix issue #235: Use torch.no_grad() in inference to prevent excessive…」(已 merged)。LiveTalking 侧对应的就是 PR #612(17.62GB→3.9GB)。

坑 6:依赖版本锁定陈旧 requirements.txt 逐字:diffusers==0.30.2、accelerate==0.28.0、numpy==1.23.5、tensorflow==2.12.0、opencv-python==4.9.0.80、transformers==4.39.2、huggingface_hub==0.30.2、librosa==0.11.0、einops==0.8.1、gradio==5.24.0。


交叉结论速览(C1 三仓库)

项目推理峰值显存(最可信数字)16GB 单卡推理训练显存实时 FPS(实测/官方)首帧延迟流式/打断中文sm_120 状态
LiveTalkingmusetalk 3.9GB(修 no_grad 后,PR#612 实测 5090);未修时 17.62GB→OOM;wav2lip ~200MB(作者估算)✅ 充足(务必含 PR#612)— (不训练,用现成权重)wav2lip256 3060 60fps;musetalk 3080Ti 42 / 3090 45 / 4090 72 [官方]未查到(WebRTC 建链 6s~数分钟,非模型延迟)✅ 支持打断、WebRTC/RTMP/多并发✅ 完整无 sm_120 issue(total 0);官方验证 torch 2.9.1+cu128
Ultralight-Digital-Human未查到实测值(模型移动端级,LiveTalking 侧估 ~1.5GB FP32 [他人估算])✅ 推断充足未查到(默认 batch=1)<10ms/帧(2080 多并发,需转 ONNX)[作者实测]未查到⚠️ 支持但代码未完善;继任 FeatherTalk 已开源 C++ 流式✅ 完整(康辉 demo)🔴 官方栈 torch1.13.1+cu117 无 sm_120 内核,必须自行迁移 cu128
MuseTalk~3.6GB(RTX 5060 Ti 16GB 实测,issue #409);OpenTalking 峰值 5.078–5.203GB;社区实测可压到 11GB 以下✅ 充足;训练不行(stage1 ≥32GB / stage2 ≥54GB,需 8×H20)官方:stage1 bs8 ≈32GB、bs32 ≈74GB;stage2 ≈54–85GB官方 30fps+ on V100;5060Ti 实测 ≈5.8fps [推算];OpenTalking 3090 28.87 / 4090 24.77TTFV 1.77s(3090)/ 2.10s(4090)[OpenTalking 第三方];首轮总 3.24–3.61s❌ 自身无打断;项目成员确认「音频的流式处理我们暂时没有研究」;打断靠 LiveTalking✅ 官方声明中/英/日🔴 torch 需 cu128(cu124/cu126 均报 no kernel image);🔴 mmcv/mmpose 在 Py3.12 装不上/编译失败,建议改用 face-alignment

针对 RTX 5060 Ti 16GB / sm_120 的三条硬结论

  1. MuseTalk 有同款硬件实测(issue #409,2026-03-18 + 2026-08-05 追加):16GB 下 ~3.6GB VRAM 跑通 361 帧、口型正确;但吞吐约 5.8fps,达不到实时,且必须用 torch 2.10.0+cu128 且绕开 mmcv(用 face-alignment)。
  2. LiveTalking 是目前唯一「原生实时 + 可打断 + 中文 + 全身」的整合方案,官方验证栈正是 CUDA 12.8(与 sm_120 对齐),且 sm_120 无已知 issue;但 musetalk 后端必须包含 2026-08-20 合并的 no_grad 修复(PR #612),否则 16GB 卡按默认 batch=16 会 OOM。想稳,用 --model wav2lip(3060 都能 60fps)或把 batch 降到 8。
  3. Ultralight-Digital-Human 是三仓库中 16GB 适配风险最高的:官方栈(torch 1.13.1 / cu117 / py3.10)在 Blackwell 上不可用,且流式代码长期未开源、推理默认跑 CPU。作者继任项目 FeatherTalk(C++ 流式、FeatherHuBERT 音频侧算力降 40×、audio+UNet <1M)更值得跟进。

补充:维护活跃度是三仓库最被低估的选型维度

截至 2026-09-22,最后提交时间分别为:LiveTalking 2026-09-13(9 天前)、Ultralight-Digital-Human 2026-07-22、MuseTalk 2025-09-26(近一年前)。 sm_120 是 2025 年后的新平台,其适配问题只能靠上游持续修。MuseTalk 上游已停滞,2026 年的 Blackwell 适配案例全部来自社区自发贡献(#409、#334)且尚未合并进主干——选择 MuseTalk 必须自行维护补丁分支。LiveTalking 在 2026 年 8–9 月仍有合并提交,是其被推荐为首选整合层的额外理由。


来源清单

官方仓库 / 文档

  1. https://github.com/lipku/LiveTalking
  2. https://raw.githubusercontent.com/lipku/LiveTalking/main/README.md
  3. https://raw.githubusercontent.com/lipku/LiveTalking/main/requirements.txt
  4. https://doc.livetalking.ai/docs/faq/
  5. https://github.com/anliyuan/Ultralight-Digital-Human
  6. https://raw.githubusercontent.com/anliyuan/Ultralight-Digital-Human/master/README.md (main 分支 README 为 404)
  7. https://github.com/anliyuan/FeatherTalk (继任项目,README 内链)
  8. https://github.com/TMElyralab/MuseTalk
  9. https://raw.githubusercontent.com/TMElyralab/MuseTalk/main/README.md
  10. https://raw.githubusercontent.com/TMElyralab/MuseTalk/main/requirements.txt
  11. https://huggingface.co/TMElyralab/MuseTalk
  12. https://arxiv.org/abs/2410.10122 (MuseTalk 技术报告)

GitHub Issues / PR(含实测数字)

  1. https://github.com/lipku/LiveTalking/pull/612 (musetalk 实时推理显存 17.6G→3.9G,2026-08-20 merged)
  2. https://github.com/lipku/LiveTalking/issues/192 (4090 musetalk OOM,17.77GiB allocated)
  3. https://github.com/lipku/LiveTalking/issues/368 (WebRTC 建链 6s~数分钟)
  4. https://github.com/lipku/LiveTalking/issues/618 (中文 ASR/TTS 后端请求,2026 活跃)
  5. https://github.com/lipku/LiveTalking/issues/535 (多并发会话)
  6. https://github.com/TMElyralab/MuseTalk/issues/409 (RTX 5060 Ti 16GB / sm_120 实测:~3.6GB VRAM,7s 音频→30s 推理;2026-03-18 及 2026-08-05 追加)
  7. https://github.com/TMElyralab/MuseTalk/issues/392 (VRAM 需求;GTX 1060 3GB 可跑但极慢)
  8. https://github.com/TMElyralab/MuseTalk/issues/310 (4090D 实时爆显存;实测可压到 11G;batch=15 约 20G)
  9. https://github.com/TMElyralab/MuseTalk/issues/334 (50 系支持;mmcv 2.0.1 不支持 cu128 → 2.1.0;5090 可跑)
  10. https://github.com/TMElyralab/MuseTalk/issues/362 (CUDA out of memory)
  11. https://github.com/TMElyralab/MuseTalk/issues/331 (OutOfMemoryError)
  12. https://github.com/TMElyralab/MuseTalk/issues/396 (社区质疑实时性)
  13. https://github.com/TMElyralab/MuseTalk/issues/248 (4090 只能 5fps)
  14. https://github.com/TMElyralab/MuseTalk/issues/49 (realtime 脚本实际比日志 fps 慢)
  15. https://github.com/TMElyralab/MuseTalk/issues/333 (推理速度问题)
  16. https://github.com/TMElyralab/MuseTalk/issues/377 (To achieve higher fps)
  17. https://github.com/TMElyralab/MuseTalk/issues/359 (比 1.0 更慢)
  18. https://github.com/TMElyralab/MuseTalk/issues/173 (Making MuseTalk 40% faster)
  19. https://github.com/TMElyralab/MuseTalk/pull/349 (推理加 torch.no_grad() 修复)
  20. https://github.com/anliyuan/Ultralight-Digital-Human/issues/120 (推理内存占用近 10G)33. https://github.com/anliyuan/Ultralight-Digital-Human/issues/74 (推理全程走 CPU,GPU 无动静)
  21. https://github.com/anliyuan/Ultralight-Digital-Human/issues/103 (催开源实时流代码)
  22. https://github.com/anliyuan/Ultralight-Digital-Human/issues/101 (音频特征提取太慢)
  23. https://github.com/anliyuan/Ultralight-Digital-Human/issues/161 (Whisper-tiny vs WeNet 速度)
  24. https://github.com/anliyuan/Ultralight-Digital-Human/issues/18 (200 epochs 效果抖)
  25. https://github.com/anliyuan/Ultralight-Digital-Human/issues/81 (移动端 wenet 特征提取位置)
  26. https://github.com/anliyuan/Ultralight-Digital-Human/issues/226 (FunASR/SenseVoice 集成建议)
  27. https://github.com/anliyuan/Ultralight-Digital-Human/issues/52 (hubert 报错)

第三方博客 / 文档 / 基准

  1. https://www.cnblogs.com/livetalking/articles/22134438 (拆解 LiveTalking 资源消耗;三模型显存估算 + FPS 表 + 40ms 预算;作者=lipku)
  2. https://datascale-ai.github.io/opentalking/latest/en/avatar_models/musetalk/ (MuseTalk 峰值显存 5.078/5.203GB、TTFV 1769/2095ms、Steady FPS 28.87/24.77)
  3. https://datascale-ai.github.io/opentalking/latest/en/benchmark/ (OpenTalking 总基准页)
  4. https://deepwiki.com/anliyuan/Ultralight-Digital-Human/3.2-digital-human-model-training (Ultralight 训练参数:epochs 200 / batchsize 1;SyncNet 参数已过时)
  5. https://hf-mirror.com/TMElyralab/MuseTalk/resolve/main/musetalkV15/unet.pth (HEAD 实测 Content-Length 3,400,074,924 B)
  6. https://hf-mirror.com/TMElyralab/MuseTalk/resolve/main/musetalk/pytorch_model.bin (HEAD 实测 Content-Length 3,400,076,549 B)

2026 年状态核实(最后提交 / Star / License)—— 无 API rate limit 的取数端点

  1. https://github.com/lipku/LiveTalking/commits/main.atom (<updated> 首条 = 2026-09-13T11:56:50Z)
  2. https://github.com/anliyuan/Ultralight-Digital-Human/commits/master.atom (<updated> 首条 = 2026-07-22T12:29:23Z;同时验证 main 分支不存在)
  3. https://github.com/TMElyralab/MuseTalk/commits/main.atom (<updated> 首条 = 2025-09-26T05:44:17Z)
  4. https://img.shields.io/github/stars/lipku/LiveTalking.json (9.6k)· https://img.shields.io/github/license/lipku/LiveTalking.json (Apache-2.0)
  5. https://img.shields.io/github/stars/anliyuan/Ultralight-Digital-Human.json (2.6k)· https://img.shields.io/github/license/anliyuan/Ultralight-Digital-Human.json (not specified)
  6. https://img.shields.io/github/stars/TMElyralab/MuseTalk.json (6.6k)· https://img.shields.io/github/license/TMElyralab/MuseTalk.json (not identifiable by github)
  7. https://img.shields.io/github/stars/anliyuan/FeatherTalk.json (93)· https://img.shields.io/github/last-commit/anliyuan/FeatherTalk.json (2026-09)

新增 issue 证据

  1. https://github.com/TMElyralab/MuseTalk/issues/33 (4090:去掉 VAE→CPU 回传可达 60+fps,实际仅约 30fps;项目成员 czk32611 确认「音频的流式处理我们暂时没有研究」)
  2. https://www.cnblogs.com/livetalking/articles/22134438 (补充:avatar 帧缓存 10s 1080p ≈ 1.61GB RAM;默认 max_session 5 路;NVENC 并发上限表;单路 1080p libx264 软编 ≈ 0.45 核;反压机制 buffer_size >= 5)

已尝试但未能取用(记录以免重复劳动)

  1. https://blog.csdn.net/lipku/article/details/148593974 (作者 lipku 的「实时数字人多并发」,CSDN 返回 2150 字节反爬壳页,正文未取到)
  2. https://blog.csdn.net/beautifulmemory/article/details/152012956 (LiveTalking Linux 搭建教程,同样被 CSDN 反爬拦截,2274 字节)
  3. https://blog.gitcode.com/3fe07b50fbd0927c431432a2e3d7d359.html (MuseTalk 4090 性能优化,页面自述由 AIGC 生成且「登录后查看全文」,不具引用价值,已弃用)
  4. https://raw.githubusercontent.com/anliyuan/FeatherTalk/main/README.md 与 /master/README.md (均返回 0 字节;FeatherTalk 的显存/延迟数字未查到,其能力描述仅能从 Ultralight README 转述)
  5. LICENSE 核验:https://raw.githubusercontent.com/lipku/LiveTalking/main/LICENSE → HTTP 200(Apache License 正文实取);.../anliyuan/Ultralight-Digital-Human/master/LICENSE 与 .../TMElyralab/MuseTalk/main/LICENSE → HTTP 000(连接失败,非 404),故二者的许可证状态未能确认,仅以 shields.io 结果为准。

检索方式说明:GitHub issue/PR 正文通过 github.com 页面内嵌 JSON(bodyHTML)解析获取。本次调研中途 api.github.com 未认证额度(core 60/hr,与同 IP 的其他 agent 共享)被耗尽,故后续改用:commits/<branch>.atom(提交时间)、shields.io(Star/License)、raw.githubusercontent.com(README)、以及 web_search + curl 目标页。所有数字均逐字取自上述来源,未做任何推测性填补;无来源处一律标注「未查到」。

下载此文件