【作者自述】=仓库 README / Release 官方口径;【第三方实测】=Issue / 博客 / Wiki 中他人给出的数字;【推断】=基于文件与文档的合理推导,非直接证据;未查到=未能验证,不编造。duixcom/Duix.Heygem 这个仓库在 2026 年已经不存在了 —— 它被官方重命名为 duixcom/Duix-Avatar,GitHub 对旧地址做 301 跳转。所以"Duix.Heygem 2026 状态"的正确答案是:主体已改名,改名后的 Duix-Avatar 是唯一活跃主线。CUDA error: no kernel image is available for execution on the device,根因指向 Docker 镜像里的 PyTorch 不含 CC 12.0 内核,该 Issue 至今 Open 且零回复。| 项目 | 结果 | 证据 |
|---|---|---|
duixcom/Duix.Heygem | 301 跳转到 duixcom/Duix-Avatar | curl -sL -w '%{url_effective}' → 200 -> https://github.com/duixcom/Duix-Avatar;ungh.cc 对 duixcom/Duix.Heygem 查询返回 "id":907627874,"name":"Duix-Avatar"(同一个 repo id,证明是重命名而非新仓) |
| Stars / Forks | 15547 / 2642 | ungh.cc/repos/duixcom/Duix-Avatar,2026-09-22 读取 |
| 创建 / 最后 push | 创建 2024-12-24T03:00:26Z;最后 push 2026-04-21T07:06:36Z | 同上。→ 2026 年有提交(4 月),但到 9 月已约 5 个月无代码更新 |
| 改名证据(Release) | v1.0.5(2025-08-15)发布说明:"The original project name "HeyGem" has now been officially changed to "Duix.Avatar".";v1.0.6(2025-09-28):"It's just a name change. If the version you're using has no issues, you can continue to use it." | releases |
duixcom org 下各仓 2026 状态(逐个核查,含"是否存在")说明:任务要求的
curl https://api.github.com/orgs/duixcom/repos因本机 IP 限流(core remaining=0,reset 13:52)无法执行;改用 ungh.cc 镜像逐个精确查询 + 落地页跳转探测,结论同样可证伪。
| 仓库 | 状态(2026-09-22) | Stars | 最后 push | 说明 |
|---|---|---|---|---|
duixcom/Duix-Avatar(原 Duix.Heygem) | ✅ 存在、2026 有更新 | 15547 | 2026-04-21 | 本组主目标的正身 |
duixcom/Duix.Mobile(原 Duix.mobile) | ✅ 存在、2026 活跃 | 8252 | 2026-08-05 | 移动端/嵌入式实时交互 SDK,非本地 GPU 服务。仓描述:"on-premise deployment and <1.5 s latency" |
duixcom/Duix.Heygem.Android | ❌ 404,不存在 | — | — | 该仓从未存在或已删除;raw 的 main/master 均 404 |
duixcom/duix-skills | ⚠️ 2026-07-14 新建,但非本地模型 | 4 | 2026-08-12 | 见下方"2026 后继者判定" |
duixcom/Duix-Reface | ❌ 404(Duix.Mobile README 里仍在推荐它,但已不可访问) | — | — | 文档链接已失效 |
duixcom/Duix.mobile(旧名) | 301 → duixcom/Duix-Mobile | — | — | 同为重命名 |
"2026 后继者是否取代 Duix.Heygem" 判定:
duixcom/duix-skills(2026-07-14 新建)不构成本地部署的后继者:它是一套给 AI agent 用的云端 API skill,需要 DUIX_APP_ID/DUIX_APP_KEY/DUIX_API_KEY 凭据,走订阅套餐 + 预付费积分计费,产出是 conversation_url 网页链接或云端生成的 MP4;其 README 明确写 "Prepaid credits (consumed by task video duration)"、"[Production time] About 20 minutes to 2 hours"。→ 它把 Duix 的 lip-sync 能力 SaaS 化了,与"本地 16GB 显卡部署"是两条路。(证据:duix-skills README)CPU: 13th Gen Intel Core i5-13400F / Memory: 32GB / Graphics Card: RTX 4070,硬盘 C 盘 >100GB、D 盘 >30GB;服务器部署要下载 约 70GB 流量。(README)GPU | NVIDIA RTX 4070 (8GB VRAM) | NVIDIA RTX 4090 / RTX 5090 (16GB+ VRAM)(Minimum → Recommended),并给出调优表:8GB → max_split_size_mb 256/512、12-16GB → 512 (default)、24GB+ → 1024。 ⚠️ 该页有硬错误需警惕:它把 RTX 5090 的 "CUDA Compute" 标为 9.0(5090 实为 sm_120),说明此页为 AI 自动生成、未经人工校对;其数字仅可作参考。链接:deepwiki.com/duixcom/Duix-Avatar/6.3-hardware-and-performance-tuningguiji2025/duix.avatar),仓库不发布独立 .pt/.safetensors,故无 fp16 权重体积数字可引用。Docker Hub 我尝试拉取 tag 体积,多次超时未成功(hub.docker.com 在本机网络下不可达)。Duix.Mobile 的时延数字不能用于 Duix.Avatar —— Mobile README 写 "AI avatar response latency under 120ms (tested on Snapdragon® 8 Gen 2 SoC)",那是手机 SoC 上的端侧渲染,与 GPU 服务的视频合成是两套东西。(Duix-Mobile README)http://127.0.0.1:8383/easy/submit,进度查询 http://127.0.0.1:8383/easy/query?code=${taskCode}(GET 轮询);TTS 接口参数里甚至有 "streaming": false, // Fixed parameter 这一固定值。→ 非流式、不可打断。(README)~/duix_avatar_data/face2face:/code/data,即在源视频上重绘人脸区域。据此推断输出构图继承源视频,源视频是半身/全身就得到半身/全身;但这属于推断,非文档明示,我未查到官方对"半身/全身"的明确能力声明。证据:deploy/docker-compose-linux.ymlIssue #624,标题 "RTX 5070 (Blackwell) not supported — CUDA kernel error on Docker backend",opened on Sep 4, 2026,状态 Open,且页面零回复("Sign up for free to join this conversation")。原文逐字引用:
GPU: NVIDIA GeForce RTX 5070 Laptop GPU Compute Capability: 12.0 (Blackwell) Driver: 592.15 Error from duix-avatar-tts container:
RuntimeError: CUDA error: no kernel image is available for execution on the deviceRoot cause: The Docker images are compiled with a PyTorch version that doesn't include CUDA kernels for compute capability 12.0. PyTorch 2.7+ added Blackwell support. Request: Please rebuild and push updated Docker images (guiji2025/fish-speech-ziming,guiji2025/fun-asr,guiji2025/duix.avatar) compiled against PyTorch 2.7+ with sm_120 support.
链接:issue #624
→ 对 RTX 5060 Ti (sm_120) 的直接含义:默认三个 Docker 镜像的主线编译目标不含 sm_120,TTS 容器会直接抛 no kernel image is available。维护者截至 2026-09-22 未回应。
deploy/docker-compose-5090.yml(HTTP 200,1064 字节),镜像换成了 guiji2025/duix.avatar-5090 与 guiji2025/fish-speech-5090,仍保留 PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512 与 shm_size: '8g'。 → 【推断】50 系走单独镜像这条路径"理论上"是给 Blackwell 的,但官方只声明在 5090 上验证过;在 5060 Ti 上是否可用未查到任何证据,且 #624 的存在说明普通镜像确实没有 sm_120 内核。这是 16GB 选型里最大的未知风险点。 (证据:docker-compose-5090.yml)三个服务(docker-compose-linux.yml):
| 服务 | 镜像 | 端口 | 关键配置 |
|---|---|---|---|
| duix-avatar-tts | guiji2025/fish-speech-ziming | 18180:8080 | runtime: nvidia,NVIDIA_VISIBLE_DEVICES=0 |
| duix-avatar-asr | guiji2025/fun-asr | 10095:10095 | runtime: nvidia,privileged: true |
| duix-avatar-gen-video | guiji2025/duix.avatar | 8383:8383 | runtime: nvidia,privileged: true,shm_size: '8g',PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512 |
坑点:
runtime: nvidia,其中两个 privileged: true → 与宿主驱动强耦合;README 明确 NVIDIA 显卡与驱动是硬性前提,且 FAQ 原文 "All computing power of this project is local. The three services won't start without an NVIDIA graphics card or proper drivers."sudo nvidia-ctk runtime configure --runtime=docker(README 给了完整步骤);没有 CPU 回退路径(DeepWiki:"The system has no CPU-only fallback; an NVIDIA GPU is mandatory for all deployment modes.")shm_size: '8g' 是硬需求:DeepWiki 把 "Bus Error (Core Dumped)" 归因于 "Insufficient shared memory → Increase shm_size to 12g or 16g";1080p/60s 就建议 8GB,4K 建议 12GB+。LICENSE),上句为 README 自述口径。同时 README 明说自部署版的定位是 "Lip Sync Effect: Usable effect"(可用级),而云端 API 才是 "Stunning and higher definition effect"、且 "Iteration Speed: Slow updates, bug fixes depend on the community" —— 官方自认开源版迭代慢。| 项目 | 数值 | 证据 |
|---|---|---|
| Stars / Forks | 13219 / 2849 | ungh.cc/repos/Rudrabha/Wav2Lip |
| 创建 | 2020-08-07 | 同上 |
| 最后 push | 2025-06-22 | 同上。→ 2026 年零更新 |
| 最新 Issue 活动 | 搜索页里最新的 Issue 是 #613 "Extracting raw audio... and stuck. Load not working",opened on Jan 3, 2024;#584 opened on Nov 10, 2023 | issues 搜索 |
停更的硬证据(README 已被改成商业导流页):master 分支 README 开头第一段不再是论文说明,而是:
Commercial Version
Create your first lipsync generation in minutes. Please note, the commercial version is of a much higher quality than the old open source model! Create your API key from the Dashboard.
其后大段内容是 Sync.so 的 pip install syncsdk 调用示例,论文原文被挤到 README 尾部。(README) → 结合 2025-06-22 那次 push 的性质(把 README 换成商业 API 导流,而非代码更新),以及 Issue 追踪器自 2024 年初起基本无维护者回应,可以判定:Wav2Lip 在 2026 年已是"存档态"项目,作者精力已转移到商业产品 sync.so。
社区推荐的后继者:
rudrabha@synclabs.so / prajwal@synclabs.so 与 https://synclabs.so/)。Python 3.6、ffmpeg。(README)Wav2Lip 与 Wav2Lip + GAN 两个 checkpoint),仓库内不发布文件;我尝试检索 CSDN/Reddit 的实测数字,未获得可引用的具体数值,故不编造。可确认的只有量级判断:未查到官方或可引用的实测数字。 唯一相关线索是 Issue #584(标题 "How much time do you need to lip sync a 10 sec or 1 minute video?",opened on Nov 10, 2023,至今 Open 且无人给出有效回答),提问者原文:
I have been trying the last days with both wav2lip HD (not in auto) and retalker, and found that both are slow and very GPU consuming. I would like to know everyone of you HOW MUCH GPU do you use (what card) and HOW MUCH time does it take for you to do it? Please contribute. Because I am about to drop this technology and give up on it, maybe others peoples experiences will give me hope.
链接:issue #584 → 该 Issue 本身没有给出数字,它只能证明"用户普遍抱怨慢、且社区无人应答"。
不支持。 架构是整段批处理:python inference.py --checkpoint_path <ckpt> --face <video.mp4> --audio <an-audio-source>,结果一次性写盘为 results/result_voice.mp4。无流式接口、无 chunk 级回调、不可中途打断。(README)
--pads 调下巴、--resize_factor 降分辨率),身体/姿态完全不动。→ 能"保留"半身/全身构图,但不具备身体动作生成能力。手部/身体穿帮是该类方法的通病(README 的 Tips 里已在教用户用 --nosmooth 处理"two mouths"等伪影)。Python 3.6)→ 这是 2020 年环境,在 Ubuntu 24.04 + CUDA 13.2 上基本无法照原样复现,需自行升级并重测。pip install -r requirements.txt,无 torch/CUDA 版本 pin;因此 sm_120 兼容性完全取决于你自己装的 torch(【推断】需 torch ≥ 2.7 + cu128 才有 sm_120 内核 —— 与 Duix #624 里"PyTorch 2.7+ added Blackwell support"的说法一致)。face_detection/detection/sfd/s3fd.pth,来源是 adrianbulat.com(README 自己提供了备用 SharePoint 链接,说明主链常挂);模型权重走 Google Drive。| 项目 | 数值 | 证据 |
|---|---|---|
| Stars / Forks | 6090 / 981 | ungh.cc/repos/bytedance/LatentSync;GitHub Issues 页 aria-label 亦为 "6090 users starred this repository" |
| 创建 / 最后 push | 2024-12-11 / 2025-06-20 | 同上 → 2026 年零代码更新(截至 2026-09-22 已停更约 15 个月) |
| GitHub Releases | 空({"releases":[]}) | 同上。权重走 HuggingFace,不打 GitHub Release |
| 最新版本 | 1.6(2025/06/11) | README Updates:"2025/06/11: We released LatentSync 1.6, which is trained on 512×512 resolution videos to mitigate the blurriness problem.";HF 仓 ByteDance/LatentSync-1.6 |
"2026 后继版是否改变要求" 的答案:没有 2026 新版;而 1.6 相对 1.5 把显存要求改坏了(8GB → 18GB)。
【作者自述,README 原文,最权威】
Minimum VRAM for inference:
- 8 GB with LatentSync 1.5
- 18 GB with LatentSync 1.6
→ 16GB 单卡 < 官方 18GB 门槛,官方口径下 1.6 不可行。 要留在 16GB 内,官方口径只允许退回 1.5(8GB)。
训练档(同一 README,供参考):stage1.yaml 23GB;stage2.yaml 30GB;stage2_efficient.yaml 20GB(作者注:"suitable for users with consumer-grade GPUs, such as the RTX 3090");stage1_512.yaml 30GB;stage2_512.yaml 55GB。(README)
fp16/权重体积(第三方):ComfyUI 封装仓 README 列出 latentsync_unet.pt (~5GB)、stable_syncnet.pt (~1.6GB),并称 "Reduced VRAM Requirements: Optimized to run on 20GB VRAM (RTX 3090 compatible)"。(ShmuelRonen/ComfyUI-LatentSyncWrapper)⚠️ 该仓为第三方封装,且"20GB / RTX 3090"表述自相矛盾(3090 是 24GB),仅作量级参考。
🟢 决定性第三方实测:RTX 5070 Ti 16GB 真的跑起来了(2026 年 7 月) Issue #365 "TensorRT-RTX optimized inference path and reduced-step speed benchmarks",opened on Jul 3, 2026,作者 Petrus Vermaak(自述 "This work was directed, tested, and validated by Petrus Vermaak... OpenAI Codex served as the implementation and engineering agent")。测试系统:RTX 5070 Ti 16 GB test system —— 同为我们关心的显存档位与 Blackwell 世代。基准夹具原文:
Fixture: official demo video/audio, 9.68s output, 242 frames, 512 face processing, 20 steps, guidance 1.5.
| Backend | Wall seconds | Sec/output second | Change vs exact20 | Peak total VRAM |
|---|---|---|---|---|
| exact20 PyTorch eager | 397.512 | 41.065 | baseline | 15807 MB |
| optimized20 TensorRT-RTX exact | 268.160 | 27.702 | 32.54% faster | 12551 MB |
(逐字引用表格数字;链接:issue #365)
→ 【第三方实测】16GB 单卡在 512×512 / 20 步下可以跑完 LatentSync 1.6:裸 PyTorch 峰值 15807 MB(已占满 16GB 的 96%,极限)、TensorRT-RTX 优化后 12551 MB(留有余量)。官方说的 18GB 是保守门槛,实测 16GB 能压进去,但 PyTorch 裸跑基本贴着天花板。
用户侧旁证(都指向 16GB 不够):
inference_steps [20-50]、guidance_scale [1.0-3.0] 两个可调项(步数越高越慢)。optimized20 779.70s、optimized12 609.50s、optimized8 502.50s、optimized4 384.70s;作者自己警告 "optimized4 was visually plausible in this test but showed the largest objective drift, so I would not present it as equivalent quality"(降步数是"近似"模式,质量会漂)。Engine cosine vs PyTorch 0.9999949036、Full-frame SSIM mean 0.981501、SyncNet confidence 8.344(baseline 8.578)。不支持,且被 Issue 明确暴露。 架构是整段扩散去噪(20 步),没有 chunk 级流水线;Issue #326 标题即为 "How to run inference on longer video?"(issue #326)→ 长视频要自己切段处理。无流式输出、无打断机制。
🔴 最硬的坑:官方依赖锁死在 cu121,不含 sm_120 内核。 requirements.txt 逐字(raw):
torch==2.5.1
torchvision==0.20.1
--extra-index-url https://download.pytorch.org/whl/cu121
diffusers==0.32.2
transformers==4.48.0
decord==0.6.0
accelerate==0.26.1
einops==0.7.0
omegaconf==2.3.0
opencv-python==4.9.0.80
mediapipe==0.10.11
python_speech_features==0.6
librosa==0.10.1
scenedetect==0.6.1
ffmpeg-python==0.2.0
imageio==2.31.1
imageio-ffmpeg==0.5.1
lpips==0.1.4
face-alignment==1.4.1
gradio==5.24.0
huggingface-hub==0.30.2
numpy==1.26.4
kornia==0.8.0
insightface==0.7.3
onnxruntime-gpu==1.21.0
DeepCache==0.1.1
→ torch==2.5.1 + cu121 官方 wheel 不含 sm_120(Blackwell)内核。在 RTX 5060 Ti 上照原样装,会撞上与 Duix #624 同型的 no kernel image is available for execution on the device。必须自行升级到 torch ≥ 2.7 + cu128/cu13x,而任何 torch 升级都可能打破下面这批老旧 pin。这是本项目在 sm_120 上最大的工作量来源。
✅ 好消息:不需要 xformers,也不需要 flash-attn。 我逐条核对了上面的 requirements:没有 xformers,也没有 flash-attn。→ 不存在 flash-attn/triton 需要现场编译 sm_120 内核的问题,这一点比很多扩散类 talking-head 项目(如 MuseTalk / 部分 ComfyUI 节点)友好得多。注意力走 diffusers 默认路径。 (若走 TensorRT-RTX 路线,Issue #365 已证明在 16GB Blackwell 上可行且比 PyTorch 省 3256MB 峰值显存、快 32.54%,是本项目在 16GB 卡上最值得走的优化路径。)
其他版本坑(逐条):
onnxruntime-gpu==1.21.0:需要与 CUDA 版本匹配的 ORT 构建;这类 pin 在 CUDA 13.x 宿主上常需换轮子。insightface==0.7.3 + face-alignment==1.4.1:人脸检测依赖,要现场编译(insightface 需 C++ 工具链),是环境搭建的常见失败点。numpy==1.26.4 + gradio==5.24.0 + mediapipe==0.10.11:整体是 2025 年上半年的老矩阵,与 torch≥2.7 的组合未被作者验证过(作者 2025-06-20 后就停更了)。DeepCache==0.1.1:推理加速用,可能与新版 diffusers 冲突(diffusers 被 pin 在 0.32.2)。setup_env.sh:README 要求 source setup_env.sh 来装依赖+下权重,但我抓取该文件返回空(未核实其内容) —— 不排除 404 或抓取失败,此项标注为未核实。能跑,但要打破官方依赖并接受"离线批处理"定位:官方 18GB 门槛把 1.6 判为不合规;实测 16GB(5070 Ti)在 512×512/20 步下峰值 15807MB 可通过,TensorRT-RTX 降到 12551MB。代价是速度仅约实时的 1/28~1/41,且必须自己把 torch 从 cu121 迁到 cu128+ 才能吃 sm_120。 它不是"实时数字人",是"离线 lip-sync 渲染器"。
| 维度 | Duix.Avatar(原 Duix.Heygem) | Wav2Lip | LatentSync |
|---|---|---|---|
| 2026 状态 | 改名后仍活跃(2026-04 有提交) | 停更(2025-06-22 最后一次 push,README 已商业导流) | 停更 15 个月(2025-06-20) |
| Stars | 15547 | 13219 | 6090 |
| 官方推理显存 | 未给数字(整机推荐 RTX 4070,16GB+ 为推荐档) | 未给 | 8GB(1.5) / 18GB(1.6) |
| 16GB 单卡 | ⚠️ 怕是不行:三容器常驻,无 16GB 实测 | ✅ 几无悬念可跑(推断) | ⚠️ 官方不合规,实测 15807MB 压线可跑 |
| 成片速度 | 未查到 FPS(官方定义非实时) | 未查到(Issue 普遍抱怨慢) | 27.7~41.1 秒 / 1 秒成片(5070Ti 16GB 实测) |
| 流式/打断 | ❌ 提交+轮询 | ❌ 整段批处理 | ❌ 整段扩散 |
| 中文 | ✅ 8 语言含中文 | ⚠️ 语言无关,但只训练英文(LRS2) | ✅ 明确优化中文视频 |
| 半身/全身 | 【推断】继承源视频(face2face) | 只改嘴部,身体不动 | ❌ 仅面部 |
| sm_120 现状 | 🔴 坏(Issue #624 Open 零回复,主线镜像无 CC12.0 内核;仅 5090 专用镜像存在,未验证 5060 Ti) | ⚪ 取决于自装 torch(无 pin) | 🟡 需自行把 cu121→cu128+(无 xformers/flash-attn,省一大坑) |
| 商用许可 | ✅ 免费商用(>10万用户或>1000万美元营收需签约) | ❌ 严格禁止商用 | 未核实(论文/ByteDance 开源,未逐字核实 LICENSE) |
| 部署形态 | Docker 三容器 + 客户端 App | 裸 Python 脚本 | 裸 Python 脚本 |
Duix 系
no kernel image is available(2026-09-04 开,Open 零回复)Wav2Lip
LatentSync
未查到 / 未能核实的项(诚实标注)
hub.docker.com 本机多次超时) —— 未查到duixcom org 完整仓库列表 —— 官方 API 限流(core remaining=0,reset 13:52),改用 ungh 逐仓探测,可能遗漏 org 下其他未探测的仓