# 实时数字人（Talking Head / Avatar）开源项目核实 —— Cluster B

**核实日期：2026-09-22**（本文件所有"last commit"均为当日实测值）

**核实方法**（全部为可复现的一手探测，未使用 2024/2025 二手描述）：

- 仓库存在性：`curl -s -o /dev/null -w "%{http_code}" https://github.com/<owner>/<repo>`（404 = 不存在）
- stars / license：`https://img.shields.io/github/stars|license|last-commit/<owner>/<repo>.json`
- **精确最后提交时间**：GitHub commits Atom feed `https://github.com/<owner>/<repo>/commits/<branch>.atom` 的 `<updated>` 字段（比 shields 的"月份"粒度更硬）
- 许可证实体：`https://raw.githubusercontent.com/<owner>/<repo>/<branch>/LICENSE`
- 模型版本存在性：`https://hf-mirror.com/api/models?search=<关键词>`
- 无法验证的一律写 **未查到**。

> ⚠️ shields.io 的 `last-commit` 只给"月份"粒度：**不带年份的裸月份 = 2026 年**（如 `march` = 2026-03），带年份才是往年。本文件已用 Atom feed 交叉核对。

---

## 1. Realtime-Venus

| 项目 | 内容 |
|---|---|
| 名称 | Realtime-Venus（Realtime-Venus-Omni · 9B / Realtime-Venus-Audio · 9B / Realtime-Venus-Harness） |
| 仓库 URL | https://github.com/inclusionAI/Realtime-Venus （✅ 存在，HTTP 200） |
| 是否存在 | **存在**（Ant Group「Venus Team」+ 清华大学，inclusionAI 组织） |
| stars | **40** |
| license | **Apache-2.0**（README「License」节明示；`LICENSE` 文件实测为 Apache License 2.0 正文） |
| last commit | **2026-09-21**（Atom feed `<updated>2026-09-21T15:45:17Z`；shields 显示 `yesterday`） |
| 模型权重 | Hugging Face [inclusionAI/Realtime-Venus](https://huggingface.co/inclusionAI/Realtime-Venus)（hf-mirror API 实测 200）；ModelScope 镜像 `inclusionAI/Realtime-Venus` |
| Release | tag `android-beta-0918`，含 `Realtime-Venus-0918.apk` Android Beta 演示包 |
| 论文 | [arXiv 2609.13814](https://arxiv.org/abs/2609.13814) |
| 简介 | **全双工（full-duplex）实时交互系统 + 异步委派（asynchronous delegation）**，不是"音频驱动人脸视频生成"模型。Omni-9B 吃**流式音视频输入**，边听边说、可主动发起回应、原生生成语音；Harness 运行时把自然语言任务丢到后台（demo 用 Codex 作为任务后端）执行，结果回到同一会话由模型择时用语音播报。仓库含 Omni 模型集成、可复用 Harness 包、浏览器 demo、Android Beta。 |

**定位提醒**：它是"实时多模态对话 + 语音输出"的交互层，**不产出 talking-head 视频**。若选型目标是"音频→口型视频"，本项只对应"实时对话脑 + TTS 语音"，需另外叠加渲染层。

来源：[GitHub](https://github.com/inclusionAI/Realtime-Venus) · [Hugging Face](https://huggingface.co/inclusionAI/Realtime-Venus) · [arXiv](https://arxiv.org/abs/2609.13814) · [项目页](https://realtime-venus.github.io/) · [Release tag](https://github.com/inclusionAI/Realtime-Venus/releases/tag/android-beta-0918)

---

## 2. Wan-Dancer / Wan2.2-S2V

### 2a. Wan-Dancer —— ✅ 真实存在，但**不是数字人/talking head**

| 项目 | 内容 |
|---|---|
| 名称 | Wan-Dancer（💃 Wan-Dancer） |
| 仓库 URL | https://github.com/Wan-Video/Wan-Dancer （✅ HTTP 200） |
| 是否存在 | **存在** |
| stars | **432** |
| license | **Apache-2.0**（README「📜 License」节明示，`LICENSE` 实测为 Apache 2.0 正文） |
| last commit | **2026-07-17**（Atom feed `<updated>2026-07-17T08:49:55Z`；shields `july`） |
| 模型权重 | [Wan-AI/Wan-Dancer-14B](https://huggingface.co/Wan-AI/Wan-Dancer-14B)；ModelScope `Wan-AI/Wan-Dancer-14B` |
| 论文 | [arXiv 2607.09581](https://arxiv.org/abs/2607.09581) |
| 简介 | **音乐→长时舞蹈视频**生成（hierarchical：全局关键帧规划 + 局部时序精修）。时间映射 RoPE 动态帧率适配、光流连续性损失、运动速度控制；720p/30fps **超过 1 分钟**，跨 5 种舞种，音频 + 文本条件。训练硬件为 8×A800 80GB。**与"音频驱动肖像/口型"无关**，属于"全身舞蹈视频生成"。 |

> ⚠️ 候选名 "Wan-Dancer" 容易与"数字人"混为一谈，实测它是**跳舞视频**，不驱动人脸口型。

### 2b. Wan2.2-S2V —— ✅ 存在，在 Wan2.2 主仓内（无独立仓库）

| 项目 | 内容 |
|---|---|
| 名称 | Wan2.2-S2V-14B（Speech-to-Video） |
| 仓库 URL | https://github.com/Wan-Video/Wan2.2 （S2V 为其中一项任务，**无** `Wan-Video/Wan2.2-S2V` 独立仓库） |
| 是否存在 | **存在**（作为 Wan2.2 的子能力） |
| stars | **18k**（整个 Wan2.2 仓库） |
| license | **Apache-2.0**（shields 实测） |
| last commit | **2026-09-21**（Atom feed `<updated>2026-09-21T06:16:22Z`；shields `yesterday`） |
| 模型权重 | [Wan-AI/Wan2.2-S2V-14B](https://huggingface.co/Wan-AI/Wan2.2-S2V-14B)（hf-mirror API 实测 200）；ModelScope `Wan-AI/Wan2.2-S2V-14B` |
| 简介 | **音频驱动的电影级视频生成**（speech-to-video），支持 480P/720P。首次发布于 **2025-08-26**（Wan2.2 的 News 节），随 Wan2.2 仓库持续维护（仓库 2026-09 仍活跃）。Wan2.2 本身是 MoE 视频基座（T2V-A14B / I2V-A14B / TI2V-5B / S2V-14B / Animate-14B）。 |
| 2026 活动 | 主仓 2026-09-21 有提交；但 **S2V 子任务本身在 README 的 Latest News 中最后更新为 2025-09-05**，2026 无 S2V 专属新版本公告。 |

> 注意区分：**Wan2.2-Animate-14B**（角色动画/替换，2025-09-19 发布）≠ **Wan2.2-S2V-14B**（音频驱动说话视频）。

来源：[Wan-Dancer GitHub](https://github.com/Wan-Video/Wan-Dancer) · [Wan-Dancer HF](https://huggingface.co/Wan-AI/Wan-Dancer-14B) · [Wan-Dancer arXiv](https://arxiv.org/abs/2607.09581) · [Wan2.2 GitHub](https://github.com/Wan-Video/Wan2.2) · [Wan2.2-S2V-14B HF](https://huggingface.co/Wan-AI/Wan2.2-S2V-14B) · [S2V 项目页](https://humanaigc.github.io/wan-s2v-webpage)

---

## 3. Hallo3 / Hallo4（fudan-generative-vision）

### 3a. Hallo3 —— ✅ 存在，但**已停更**

| 项目 | 内容 |
|---|---|
| 名称 | Hallo3（Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer） |
| 仓库 URL | https://github.com/fudan-generative-vision/hallo3 （✅ HTTP 200） |
| 是否存在 | **存在** |
| stars | **1.4k** |
| license | **MIT**（shields 实测 `MIT`） |
| last commit | **2025-03-13**（Atom feed `<updated>2025-03-13T08:34:23Z`；shields `march 2025`） |
| 2026 活动 | **无**（最后提交为止于 2025-03） |
| 简介 | 复旦 generative-vision 组的肖像动画（video diffusion transformer 驱动），是 Hallo 系列第三代。 |

### 3b. Hallo4 —— ✅ 存在，star 极少，模型权重**门控**

| 项目 | 内容 |
|---|---|
| 名称 | Hallo4（Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization） |
| 仓库 URL | https://github.com/fudan-generative-vision/hallo4 （✅ HTTP 200，**默认分支 master**，仅 5 个 commit） |
| 是否存在 | **存在** |
| stars | **38**（fork 4，issues 2，PR 1 —— 页面实测） |
| license | **未指定**（shields 返回 `not specified`；仓库文件列表仅 `.gitignore` / `README.md` / `inf.sh` / `requirements.txt` / `assets/` / `vace/`，**无 LICENSE 文件**） |
| last commit | **2025-11-30**（Atom feed `<updated>2025-11-30T06:26:14Z`；shields `november 2025`） |
| 2026 活动 | **无**（最后提交止于 2025-11-30） |
| 模型权重 | Hugging Face [fudan-generative-ai/hallo4](https://huggingface.co/fudan-generative-ai/hallo4) —— **门控（gated）**：匿名访问返回 **HTTP 401** `Access to model ... is restricted. You must have access to it and be authenticated` |
| 简介 | SIGGRAPH Asia 2025 论文（[ACM DL](https://dl.acm.org/doi/full/10.1145/3757377.3763914) / [ar5iv 2505.23525](https://ar5iv.labs.arxiv.org/html/2505.23525)）。用 **DPO（Direct Preference Optimization）+ 时序运动调制**做高保真动态肖像动画。作者：Jiahao Cui, Baoyou Chen, Mingwang Xu, Hanlin Shang, Yuxuan Chen, Yun Zhan, Zilong Dong, Yao Yao, Jingdong Wang, Siyu Zhu（复旦 / 百度 / 上海创智学院 / 南大 / 阿里）。推理要求 Ubuntu 20.04/22.04 + CUDA 12.1，**测试显卡为 H100**，依赖 Wan2.1 编码器。 |

> ⚠️ **Hallo4 实用性问题**：38 star、5 commit、无 LICENSE、权重门控、需 H100 级显存 —— 作为选型候选**工程可用性很低**，且 2026 年零活动。

来源：[hallo3 GitHub](https://github.com/fudan-generative-vision/hallo3) · [hallo4 GitHub](https://github.com/fudan-generative-vision/hallo4) · [hallo4 HF（门控）](https://huggingface.co/fudan-generative-ai/hallo4) · [ACM DL](https://dl.acm.org/doi/full/10.1145/3757377.3763914) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2505.23525)

---

## 4. EchoMimic V3（antgroup/echomimic_v3）

| 项目 | 内容 |
|---|---|
| 名称 | EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation |
| 仓库 URL | https://github.com/antgroup/echomimic_v3 （✅ HTTP 200） |
| 是否存在 | **存在** |
| stars | **1.1k** |
| license | **Apache-2.0**（`LICENSE.txt` 实测为 Apache License 2.0 正文；注意文件名是 `LICENSE.txt`，`LICENSE` 返回 404） |
| last commit | **2026-03-18**（Atom feed `<updated>2026-03-18T03:31:56Z`；shields `march` = 2026-03） |
| 2026 活动 | **有** |
| 模型权重 | [BadToBest/EchoMimicV3](https://huggingface.co/BadToBest/EchoMimicV3)（hf-mirror API 实测 200）；ModelScope `BadToBest/EchoMimicV3` |
| 论文 | [arXiv 2507.03905](https://arxiv.org/abs/2507.03905)；被 **AAAI 2026** 接收（2025-11-09） |
| 简介 | 统一多模态多任务人体动画：**1.3B 参数**，音频/姿态/文本多条件驱动。2025-08-08 开源代码 + 权重；2025-08-12 提供 GradioUI（**12GB 显存即可生成**）；被社区集成进 ComfyUI（[ComfyUI_EchoMimic](https://github.com/smthemex/ComfyUI_EchoMimic)，16GB 显存可跑）。 |

### 🔥 2026 关键更新（选型重点）

- **[2026.01.22]** 发布 **EchoMimicV3-Flash**（`echomimicv3-flash-pro`）到 Hugging Face：[BadToBest/EchoMimicV3/tree/main/echomimicv3-flash-pro](https://huggingface.co/BadToBest/EchoMimicV3/tree/main/echomimicv3-flash-pro) —— 这是该仓库 2026 年的实质新增能力（加速版）。
- 仓库 2026-03-18 仍有提交，是本清单里**少数 2026 年活跃且低显存门槛（12GB）**的中文团队数字人项目。

来源：[GitHub](https://github.com/antgroup/echomimic_v3) · [HF 模型](https://huggingface.co/BadToBest/EchoMimicV3) · [EchoMimicV3-Flash](https://huggingface.co/BadToBest/EchoMimicV3/tree/main/echomimicv3-flash-pro) · [arXiv](https://arxiv.org/abs/2507.03905) · [项目页](https://antgroup.github.io/ai/echomimic_v3/)

---

## 5. Hunyuan-Avatar / HunyuanVideo-Avatar（腾讯混元）

| 项目 | 内容 |
|---|---|
| 名称 | HunyuanVideo-Avatar（**仓库名就是 HunyuanVideo-Avatar**；未找到单独的 "Hunyuan-Avatar" 仓库） |
| 仓库 URL | https://github.com/Tencent-Hunyuan/HunyuanVideo-Avatar （✅ HTTP 200） |
| 是否存在 | **存在**（"Hunyuan-Avatar" 这一写法 **未找到** 对应仓库） |
| stars | **2.2k** |
| license | **Tencent Hunyuan Community License Agreement**（`LICENSE` 实测正文首行 `TENCENT HUNYUAN COMMUNITY LICENSE AGREEMENT`，Release Date: May…；shields 返回 `not identifiable by github`，因为它是自定义社区许可而非 SPDX 标准协议） |
| last commit | **2025-12-16**（Atom feed `<updated>2025-12-16T12:32:46Z`；shields `december 2025`） |
| 2026 活动 | **无**（最后提交止于 2025-12；README News 最后一条为 2025-06-06） |
| 模型权重 | [tencent/HunyuanVideo-Avatar](https://huggingface.co/tencent/HunyuanVideo-Avatar)（hf-mirror API 实测 200） |
| 简介 | 基于 **MM-DiT** 的多模态扩散 Transformer 数字人：① 角色图像注入模块（替代加性条件，消除训练/推理条件错配）；② **音频情绪模块 AEM**（从情绪参考图迁移情绪风格）；③ **人脸感知音频适配器 FAA**（latent 级人脸 mask，支持**多角色对话**独立音频注入）。2025-05-28 开源推理代码 + 权重；2025-06-06 经 [Wan2GP](https://github.com/deepbeepmeep/Wan2GP) 支持**单卡 10GB 显存**（含 TeaCache）。 |

> ⚠️ 2026 年**零提交**，作为"2026 实时数字人"候选属于**停更状态**。

来源：[GitHub](https://github.com/Tencent-Hunyuan/HunyuanVideo-Avatar) · [HF](https://huggingface.co/tencent/HunyuanVideo-Avatar)

---

## 6. OmniHuman-1.5（字节跳动）

| 项目 | 内容 |
|---|---|
| 名称 | OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation |
| 仓库 URL | **未找到**（无任何官方公开 GitHub 仓库） |
| 是否存在 | **非开源**。实测以下路径全部 **HTTP 404**：`bytedance/OmniHuman`、`bytedance/OmniHuman-1`、`bytedance/OmniHuman-1.5`、`ByteDance/OmniHuman-1.5`、`bytedance/omnihuman` |
| stars | 未查到（无仓库） |
| license | 未查到（无开源许可） |
| last commit | 未查到 |
| Hugging Face 权重 | **未查到**官方权重。hf-mirror 搜索 `OmniHuman` 仅返回社区自建评测集 [`julia527/omnihuman_benchmark`](https://huggingface.co/julia527/omnihuman_benchmark)，**非** ByteDance 官方模型 |
| 论文 | [arXiv 2508.19209](https://ar5iv.labs.arxiv.org/html/2508.19209)（Intelligent Creation Lab, ByteDance） |
| **结论** | ✅ **API-only / 闭源**。仅通过第三方推理平台提供：fal.ai 端点 [`fal-ai/bytedance/omnihuman/v1.5`](https://fal.ai/models/fal-ai/bytedance/omnihuman/v1.5)（HTTP 200）与 [Replicate `bytedance/omni-human-1.5`](https://replicate.com/bytedance/omni-human-1.5)（HTTP 200）。Replicate 描述为"A film-grade digital human model that generates realistic video from a single image, audio clip, and optional text prompt"。 |
| 简介 | 字节跳动"认知模拟"数字人：单图 + 音频（+ 可选文本提示）→ 逼真说话视频，号称具备"主动思维"（逻辑推理驱动动作）。 |

> **选型含义**：OmniHuman-1.5 **不能自部署、不可审计、按调用付费**。若硬性要求开源可自托管，本项直接排除。

来源：[fal.ai API 端点](https://fal.ai/models/fal-ai/bytedance/omnihuman/v1.5) · [Replicate](https://replicate.com/bytedance/omni-human-1.5) · [arXiv 2508.19209](https://ar5iv.labs.arxiv.org/html/2508.19209)

---

## 7. FantasyTalking2（Fantasy-AMAP）

| 项目 | 内容 |
|---|---|
| 名称 | FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation |
| 仓库 URL | https://github.com/Fantasy-AMAP/fantasy-talking2 （✅ HTTP 200，**注意实际仓库名是全小写 `fantasy-talking2`**，`FantasyTalking2` 拼写返回 404） |
| 是否存在 | **存在**（但**仅论文页/摘要**，未见推理代码或权重） |
| stars | **65** |
| license | **未指定**（shields `not specified`；`LICENSE` 与 `LICENSE` on master 均 404 —— 无许可证文件） |
| last commit | **2025-08-18**（Atom feed `<updated>2025-08-18T06:52:07Z`；shields `august 2025`） |
| 2026 活动 | **无**（最后提交止于 2025-08-18） |
| 简介 | 音频驱动肖像动画的**偏好对齐**工作：① **Talking-Critic** 多模态奖励模型量化多维人类偏好；② **Talking-NSQ** 大规模多维偏好数据集（**410K** 偏好对）；③ **TLPO**（Timestep-Layer adaptive multi-expert Preference Optimization），把偏好解耦为专家模块并按 timestep/网络层融合。声称在口型准确度、运动自然度、视觉质量上均优于基线。 |
| 论文 | [arXiv 2508.11255](https://arxiv.org/abs/2508.11255)（AAAI 2026，见仓库标题）· [项目页](https://fantasy-amap.github.io/fantasy-talking2/) |
| 代码可用性 | 实测 README 全文仅含 Abstract + Citation，**未找到 inference 代码 / 权重 / 安装说明**。README 中无 HF 权重链接。 |

> ⚠️ **名称陷阱**：任务给的 "FantasyTalking2" 大小写**不是**真实仓库名；真实 URL 是 `Fantasy-AMAP/fantasy-talking2`。

来源：[GitHub](https://github.com/Fantasy-AMAP/fantasy-talking2) · [arXiv](https://arxiv.org/abs/2508.11255) · [项目页](https://fantasy-amap.github.io/fantasy-talking2/)

---

## 8. Ditto（antgroup/ditto-talkinghead）

| 项目 | 内容 |
|---|---|
| 名称 | Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis |
| 仓库 URL | https://github.com/antgroup/ditto-talkinghead （✅ HTTP 200；`digital-human/ditto-talkinghead` 与 `di74/Ditto` 均 404） |
| 是否存在 | **存在**（蚂蚁集团 Ant Group） |
| stars | **891** |
| license | **Apache-2.0**（README「⚖️ License」节明示；`LICENSE` 实测为 Apache 2.0 正文） |
| last commit | **2025-11-12**（Atom feed `<updated>2025-11-12T13:58:42Z`；shields `november 2025`） |
| 2026 活动 | **无**（最后提交止于 2025-11-12） |
| 模型权重 | [digital-avatar/ditto-talkinghead](https://huggingface.co/digital-avatar/ditto-talkinghead) |
| 论文 | [arXiv 2411.19509](https://arxiv.org/abs/2411.19509)，被 **ACM MM 2025** 接收（2025-07-07） |
| 简介 | **运动空间扩散**的可控**实时**说话头合成。Updates：2025-01-10 开源推理代码 + 模型；2025-01-21 更新 Colab demo；2025-07-11 发布 PyTorch 模型；**2025-11-12 开源训练代码**（`tree/train` 分支，作者说明"版本较多、整理时间有限，可能与论文版略有差异"）。 |

来源：[GitHub](https://github.com/antgroup/ditto-talkinghead) · [HF](https://huggingface.co/digital-avatar/ditto-talkinghead) · [arXiv](https://arxiv.org/abs/2411.19509) · [项目页](https://digital-avatar.github.io/ai/Ditto/)

---

## 9. MEMO（MEMO: Memory-Guided Diffusion）

| 项目 | 内容 |
|---|---|
| 名称 | MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation |
| 仓库 URL | https://github.com/memoavatar/MEMO （✅ HTTP 200；`MEMO-Avatar/MEMO`、`magic-research/memo`、`antonioo/MEMO` 均 404） |
| 是否存在 | **存在** |
| stars | **1.1k** |
| license | **Apache-2.0**（shields 实测；`LICENSE` 实拉为 Apache License 2.0 正文） |
| last commit | **2025-08-06**（Atom feed `<updated>2025-08-06T15:27:15Z`；shields `august 2025`） |
| 2026 活动 | **无**（最后提交止于 2025-08-06） |
| 模型权重 | [memoavatar/memo](https://huggingface.co/memoavatar/memo)（hf-mirror API 实测 200） |
| 论文 | [arXiv 2412.04448](https://arxiv.org/abs/2412.04448) · [项目页](https://memoavatar.github.io) |
| 简介 | **记忆引导扩散**的表情丰富说话视频生成。作者：Longtao Zheng, Yifan Zhang, Hanzhong Guo, Jiachun Pan, Zhenxiong Tan, Jiahao Lu, Chuanxin Tang, Bo An, Shuicheng Yan（NTU 等）。 |
| ⚠️ 重要限制 | README 明写：仓库只包含 **MEMO-preview 模型**的示例推理脚本；「To reduce potential risks, we have only open-sourced a **preview model** for research purposes.」——**仅开源预览模型，非完整版**。生态：社区 [ComfyUI 集成](https://github.com/if-ai/ComfyUI-IF_MemoAvatar/tree/main)、[Gradio app](https://github.com/camenduru/memo-tost/blob/main/worker_runpod_gradio.py)、[Jupyter notebook](https://github.com/camenduru/memo-jupyter)。 |

来源：[GitHub](https://github.com/memoavatar/MEMO) · [HF](https://huggingface.co/memoavatar/memo) · [arXiv](https://arxiv.org/abs/2412.04448) · [项目页](https://memoavatar.github.io)

---

## 10. Sonic（tencent-ailab Sonic）

| 项目 | 内容 |
|---|---|
| 名称 | Sonic: Shifting Focus to Global Audio Perception in Portrait Animation |
| 仓库 URL | https://github.com/jixiaozhong/Sonic （✅ HTTP 200）—— ⚠️ **`tencent-ailab/Sonic` 不存在（404）**，真实托管在作者个人账号 `jixiaozhong` 下 |
| 是否存在 | **存在** |
| stars | **3.3k** |
| license | **CC BY-NC-SA 4.0**（`LICENSE` 实拉正文首行 `Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International`）—— **非商用** |
| last commit | **2026-01-08**（Atom feed `<updated>2026-01-08T11:32:14Z`；shields `january` = 2026-01） |
| 2026 活动 | **有**（截至 2026-01-08） |
| 论文 | CVPR 2025（[CVPR 开放获取 PDF](https://openaccess.thecvf.com/content/CVPR2025/papers/Ji_Sonic_Shifting_Focus_to_Global_Audio_Perception_in_Portrait_Animation_CVPR_2025_paper.pdf)）· [项目页](https://jixiaozhong.github.io/Sonic/) |
| 简介 | 强调**全局音频感知**的肖像动画（CVPR 2025）。生态：官方 Gradio demo、HF Space `xiaozhongji/Sonic`、社区 [ComfyUI_Sonic](https://github.com/smthemex/ComfyUI_Sonic)；团队 2025-05-06 另开源情绪表达系统 [DICE-Talk](https://github.com/toto222/DICE-Talk)。 |

> ⚠️ **商用限制**：CC BY-NC-SA 4.0（署名-非商业-相同方式共享），README 明确提示商业化需另行处理。

来源：[GitHub（jixiaozhong/Sonic）](https://github.com/jixiaozhong/Sonic) · [项目页](https://jixiaozhong.github.io/Sonic/) · [CVPR 2025 论文](https://openaccess.thecvf.com/content/CVPR2025/papers/Ji_Sonic_Shifting_Focus_to_Global_Audio_Perception_in_Portrait_Animation_CVPR_2025_paper.pdf)

---

## 11. FLOAT（deepbrainai-research/float）

| 项目 | 内容 |
|---|---|
| 名称 | FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait |
| 仓库 URL | https://github.com/deepbrainai-research/float （✅ HTTP 200） |
| 是否存在 | **存在**（DeepBrain AI Research） |
| stars | **492** |
| license | **CC BY-NC-ND 4.0**（README「❗License❗」节 + 2025-02-17 Updates 均指向 [creativecommons.org/licenses/by-nc-nd/4.0](https://creativecommons.org/licenses/by-nc-nd/4.0/)；仓库内**无 LICENSE 文件**，`/LICENSE` 返回 404，故 shields 显示 `not identifiable by github`）—— **非商用 + 禁止演绎** |
| last commit | **2025-11-10**（Atom feed `<updated>2025-11-10T14:06:06Z`；shields `november 2025`） |
| 2026 活动 | **无**（最后提交止于 2025-11-10） |
| 论文 | [arXiv 2412.01064](https://arxiv.org/abs/2412.01064)，被 **ICCV 2025** 接收（2025-06-26） |
| 简介 | 基于 **flow matching** 的音频驱动说话肖像生成。核心：不用像素级 latent，而用**学习到的正交运动 latent 空间**（orthogonal motion latent space），实现时序一致的 motion 生成与编辑；transformer 向量场预测器 + 逐帧条件机制；支持**语音驱动情绪增强**（speech-driven emotion enhancement）。作者：Taekyung Ki, Dongchan Min, Gyeongsu Chae。2025-02-17 开源推理代码 + checkpoints。 |

来源：[GitHub](https://github.com/deepbrainai-research/float) · [arXiv](https://arxiv.org/abs/2412.01064) · [项目页](https://deepbrainai-research.github.io/float/)

---

## 12. MuseTalk 1.5 / 2.0（TMElyralab/MuseTalk）

| 项目 | 内容 |
|---|---|
| 名称 | MuseTalk（Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling） |
| 仓库 URL | https://github.com/TMElyralab/MuseTalk （✅ HTTP 200） |
| 是否存在 | **存在** |
| stars | **6.6k** |
| license | **MIT**（`master/LICENSE` 实拉正文首行 `MIT License`；`main/LICENSE` 返回 000/超时，shields 因默认分支探测失败显示 `not identifiable by github` —— 以实体文件为准，**MIT**） |
| last commit | **2025-09-26**（Atom feed `<updated>2025-09-26T05:44:17Z`；shields `september 2025`） |
| 2026 活动 | **无**（最后提交止于 2025-09-26） |

### 版本核实结论：**2026 年最新公开版本仍是 1.5，不存在 MuseTalk 2.0**

- README「🔥 Updates」与「News」节：**[03/28/2025] 发布 1.5 版本**（相对 1.0 在清晰度、身份一致性、口型-语音同步上显著提升）；此前 **[10/18/2024]** 发布技术报告 v2。**README 中完全没有 2.0 的任何条目**。
- HF 模型仓库实测（hf-mirror `api/models?search=MuseTalk`）：官方只有 **`TMElyralab/MuseTalk`** 一个，**没有 MuseTalk 2.0 权重**。列表中的 `mlx-community/MuseTalk-1.5-q4`、`kevinwang676/MuseTalk1.5` 等均为社区 1.5 衍生/量化，非官方 2.0。
- 推理/训练代码与 **MuseTalk 1.5** 权重均已开放（README：「The inference codes, training codes and model weights of MuseTalk 1.5 are all available now!」）；支持 `v1.5` 与 `v1` 双版本推理，含 **realtime** 推理脚本（`scripts/realtime_inference.py`，示例 `--fps 25`）。
- 技术报告：[arXiv 2410.10122](https://arxiv.org/abs/2410.10122)（v3 标题为 "MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling"）。

| 模型权重 | [TMElyralab/MuseTalk](https://huggingface.co/TMElyralab/MuseTalk)（hf-mirror API 实测 200） |
|---|---|
| 简介 | 实时唇形同步（video dubbing / lip-sync）。1.5 版引入 perceptual loss + GAN loss + sync loss 联合训练，两阶段训练策略 + 时空数据采样，在视觉质量与口型精度间取平衡。**是"音改口型"而非"整脸动画"**，因此常与 Wav2Lip/LatentSync 同类对比。 |

> **"MuseTalk 2.0" = 未找到**（无仓库 tag、无 HF 权重、无 README 条目、无 release）。

来源：[GitHub](https://github.com/TMElyralab/MuseTalk) · [HF](https://huggingface.co/TMElyralab/MuseTalk) · [arXiv 2410.10122](https://arxiv.org/abs/2410.10122)

---

## 13. LatentSync 1.6 / 2.x（bytedance/LatentSync）

| 项目 | 内容 |
|---|---|
| 名称 | LatentSync（End-to-end lip sync leveraging audio conditioned latent diffusion models） |
| 仓库 URL | https://github.com/bytedance/LatentSync （✅ HTTP 200） |
| 是否存在 | **存在** |
| stars | **6.1k** |
| license | **Apache-2.0**（shields 实测 `Apache-2.0`） |
| last commit | **2025-06-20**（Atom feed `<updated>2025-06-20T07:36:51Z`；shields `june 2025`） |
| 2026 活动 | **无**（最后提交止于 2025-06-20） |

### 版本核实结论：**最新为 1.6，不存在 2.x**

- README「🔥 Updates」节列出的全部版本：
  - **`2025/06/11`：发布 LatentSync 1.6** —— 在 **512×512** 分辨率视频上训练，缓解模糊问题（changelog: `docs/changelog_v1.6.md`）。
  - `2025/03/14`：发布 LatentSync 1.5 —— 增加时序层（temporal layer）改善时序一致性、提升中文视频表现、把 stage2 训练显存降到 **20 GB**。
  - **无任何 2.0 / 2.x 条目**。
- HF 模型实测（hf-mirror `api/models?search=LatentSync`）：官方 ByteDance 仓库只有 **`ByteDance/LatentSync`、`ByteDance/LatentSync-1.5`、`ByteDance/LatentSync-1.6`** 三个；**没有 2.x**。列表中 `chunyu-li/LatentSync`、`ReopenAI/LatentSync_FT`、`Xuttt123/latentsync-pruned`、`khaduyen1993/latentsync` 等均为社区微调/剪枝，非官方。
- README 给出的显存要求：**LatentSync 1.6 需 18 GB**（`latentsync_unet.pt` + `tiny.pt` 从 [ByteDance/LatentSync-1.6](https://huggingface.co/ByteDance/LatentSync-1.6) 下载）。
- 论文：[arXiv 2412.09262](https://arxiv.org/abs/2412.09262)。
- 简介 | 端到端**唇形同步**：用 audio-conditioned latent diffusion 直接在 latent 空间做口型对齐，宣传为业界最强 lip-sync 精度之一。**属于"改口型"类，不是全身/全脸数字人动画**。

> **"LatentSync 2.x" / "LatentSync 2.0" = 未找到**（无 README 条目、无 HF 权重、无 release）。第三方博客提到 2026 年语境（如 sync.so 的 2026 评测文），但**不代表存在 2.x 版本**。

来源：[GitHub](https://github.com/bytedance/LatentSync) · [HF 1.6](https://huggingface.co/ByteDance/LatentSync-1.6) · [arXiv](https://arxiv.org/abs/2412.09262)

---

# 🆕 额外发现：2026 年值得注意的实时流式数字人开源项目

> 这些不在原始候选清单内，但在核实过程中确认是 **2026 年新建/活跃的实时流式（streaming）avatar** 项目，且**工程指标（FPS/延迟/消费级显卡可跑）明确公开**，比清单里多数"停更"候选更贴合"实时数字人"目标。

## A1. SoulX-FlashTalk（Soul-AILab，2026）

| 项目 | 内容 |
|---|---|
| 仓库 URL | https://github.com/Soul-AILab/SoulX-FlashTalk （✅ HTTP 200） |
| stars | **1.5k** |
| license | **Apache-2.0**（shields 实测） |
| last commit | **2026-07-30**（Atom feed `<updated>2026-07-30T09:13:58Z`；shields `july`） |
| 权重 | [Soul-AILab/SoulX-FlashTalk-14B](https://huggingface.co/Soul-AILab/SoulX-FlashTalk-14B)（hf-mirror API 实测 200） |
| 简介 | **首个做到亚秒级启动延迟（0.87s）+ 实时 32 FPS（8×H800 节点）的 14B 模型**。全称："Real-Time Infinite Streaming of Audio-Driven Avatars via Self-Correcting Bidirectional Distillation"。 |
| 时间线 | 2025-12-30 技术报告 + 项目页；**2026-01-08 开源推理代码 + 模型权重**。 |
| 论文 | [arXiv 2512.23379](https://arxiv.org/pdf/2512.23379) |

## A2. SoulX-FlashHead（Soul-AILab，2026）—— ⭐ 消费级显卡可实时

| 项目 | 内容 |
|---|---|
| 仓库 URL | https://github.com/Soul-AILab/SoulX-FlashHead （✅ HTTP 200） |
| stars | **1.1k** |
| license | **Apache-2.0**（shields 实测） |
| last commit | **2026-05**（shields `may` = 2026-05；Atom feed 当日探测超时，未能取得精确日） |
| 权重 | [Soul-AILab/SoulX-FlashHead-1_3B](https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B)（hf-mirror API 实测 200）· 数据集 [Soul-AILab/VividHead](https://huggingface.co/datasets/Soul-AILab/VividHead) |
| 简介 | **"Oracle-guided Generation of Infinite Real-time Streaming Talking Heads"**。明确给出消费级显卡指标：**Model_Lite 单张 RTX 4090 可达 96 FPS**（或 3 路并发各 25+ FPS 实时）；**Model_Pro 单张 RTX 4090 为 10.8 FPS**，两张 RTX 5090 可实时（25+ FPS）。 |
| 生态 | Gradio app（2026-03-04）、[ComfyUI 节点 ComfyUI_RH_FlashHead](https://github.com/HM-RunningHub/ComfyUI_RH_FlashHead)（2026-03-02）、HF Space 在线 demo（2026-03-09） |
| 时间线 | 2026-02-07 技术报告 + 数据集；**2026-02-12 开源推理代码 + 权重 + 项目页** |
| 论文 | [arXiv 2602.07449](https://arxiv.org/pdf/2602.07449) |

> 💡 **对本地 16GB 显存（RTX 5060 Ti）选型最有参考价值的一项**：1.3B 模型、Model_Lite 在 4090 上 96 FPS，是清单中唯一给出"消费级单卡实时"明确数字的流式说话头项目。

## A3. LiveAvatar（Alibaba-Quark，ECCV 2026 Spotlight）

| 项目 | 内容 |
|---|---|
| 仓库 URL | https://github.com/Alibaba-Quark/LiveAvatar （✅ HTTP 200） |
| stars | **2.4k** |
| license | **Apache-2.0**（README「📜 License Agreement」：主体 Apache 2.0；基座 Wan 模型同为 Apache 2.0） |
| last commit | **2026-08-24**（Atom feed `<updated>2026-08-24T15:10:10Z`；shields `august`） |
| 权重 | [Quark-Vision/Live-Avatar](https://huggingface.co/Quark-Vision/Live-Avatar)（hf-mirror API 实测 200） |
| 简介 | **算法-系统协同设计**的实时流式无限长交互 avatar：**14B 扩散模型**，多卡 H800 上 **45 FPS**、**4-step** 采样，Block-wise Autoregressive 处理支持 **10,000+ 秒**流式视频。 |
| 时间线 | 2025-12-04 论文 + demo 页；**2025-12-08 实时推理代码 + 权重**；**2025-12-12 单卡 80GB 推理代码**；**2026-01-20 v1.1**（FP8 量化 → 可用 **48GB 显卡**推理；编译 + cuDNN attention 提速峰值 ~2.5×、均值 3×，多 H800 稳定 45+ FPS）；**2026-06-18 被 ECCV 2026 接收为 Spotlight**。 |
| 论文 | [arXiv 2512.04677](https://arxiv.org/abs/2512.04677) |

> ⚠️ 显存门槛高：v1.1 的 FP8 路线仍需 48GB 卡；16GB 消费卡不可行。

## A4. Linly-Talker-Stream（Kedreamix）

| 项目 | 内容 |
|---|---|
| 仓库 URL | https://github.com/Kedreamix/Linly-Talker-Stream （✅ HTTP 200） |
| stars | **133** |
| license | **Apache-2.0**（shields 实测） |
| last commit | **2026-02**（shields `february` = 2026-02） |
| 简介 | "Linly-Talker-Stream: Real-Time Streaming Conversational Digital Human System —— 全双工、低延迟、实时交互数字人框架"。是 Linly-Talker 系列的**流式全双工**版本，偏**系统集成**（ASR + LLM + TTS + 渲染编排）而非单一生成模型。 |

## A5. SentiAvatar（SentiAvatar/SentiAvatar）—— 3D 交互数字人

| 项目 | 内容 |
|---|---|
| 仓库 URL | https://github.com/SentiAvatar/SentiAvatar （✅ HTTP 200） |
| stars | **453** |
| license | **未指定**（shields `not identifiable by github`；**未查到**明确开源许可，需自行核对仓库 LICENSE） |
| last commit | **2026-04**（shields `april` = 2026-04） |
| 简介 | SentiPulse 与 GSAI 联合开源，被称为**首个交互式 3D 数字人框架**，2026 年 4 月全球开源（[人民网报道](http://finance.people.com.cn/BIG5/n1/2026/0409/c1004-40697921.html) · [PRNewswire 稿](https://m.en.yna.co.kr/view/RPR20260409008000353?section=press-release/index)）。与上述 2D 视频生成路线不同，属 **3D 驱动**方向。 |

---

# 📊 汇总速查表

| # | 名称 | 仓库 | 存在 | stars | license | last commit | 2026 活跃 | 类型 |
|---|---|---|---|---|---|---|---|---|
| 1 | Realtime-Venus | [inclusionAI/Realtime-Venus](https://github.com/inclusionAI/Realtime-Venus) | ✅ | 40 | Apache-2.0 | 2026-09-21 | ✅ | 全双工对话/异步委派（含 TTS，非视频生成） |
| 2a | Wan-Dancer | [Wan-Video/Wan-Dancer](https://github.com/Wan-Video/Wan-Dancer) | ✅ | 432 | Apache-2.0 | 2026-07-17 | ✅ | 音乐→舞蹈视频（非数字人） |
| 2b | Wan2.2-S2V | [Wan-Video/Wan2.2](https://github.com/Wan-Video/Wan2.2) | ✅ | 18k(整仓) | Apache-2.0 | 2026-09-21 | ✅ | 音频驱动说话视频 |
| 3a | Hallo3 | [fudan-generative-vision/hallo3](https://github.com/fudan-generative-vision/hallo3) | ✅ | 1.4k | MIT | 2025-03-13 | ❌ | 肖像动画 |
| 3b | Hallo4 | [fudan-generative-vision/hallo4](https://github.com/fudan-generative-vision/hallo4) | ✅ | 38 | 未指定 | 2025-11-30 | ❌ | 肖像动画（DPO）；权重门控 |
| 4 | EchoMimic V3 | [antgroup/echomimic_v3](https://github.com/antgroup/echomimic_v3) | ✅ | 1.1k | Apache-2.0 | 2026-03-18 | ✅ | 多模态人体动画，1.3B，12GB 可跑 |
| 5 | HunyuanVideo-Avatar | [Tencent-Hunyuan/HunyuanVideo-Avatar](https://github.com/Tencent-Hunyuan/HunyuanVideo-Avatar) | ✅ | 2.2k | 腾讯混元社区许可 | 2025-12-16 | ❌ | 多角色情绪数字人 |
| 6 | OmniHuman-1.5 | **无仓库** | ❌ 闭源 | 未查到 | 未查到 | 未查到 | API-only | 闭源 API（fal / Replicate） |
| 7 | FantasyTalking2 | [Fantasy-AMAP/fantasy-talking2](https://github.com/Fantasy-AMAP/fantasy-talking2) | ✅ | 65 | 未指定 | 2025-08-18 | ❌ | 偏好对齐（仅摘要，无代码） |
| 8 | Ditto | [antgroup/ditto-talkinghead](https://github.com/antgroup/ditto-talkinghead) | ✅ | 891 | Apache-2.0 | 2025-11-12 | ❌ | 实时说话头（含训练码） |
| 9 | MEMO | [memoavatar/MEMO](https://github.com/memoavatar/MEMO) | ✅ | 1.1k | Apache-2.0 | 2025-08-06 | ❌ | 记忆引导（仅 preview 模型） |
| 10 | Sonic | [jixiaozhong/Sonic](https://github.com/jixiaozhong/Sonic) | ✅ | 3.3k | CC BY-NC-SA 4.0 | 2026-01-08 | ⚠️ 止于 1 月 | 肖像动画（非商用） |
| 11 | FLOAT | [deepbrainai-research/float](https://github.com/deepbrainai-research/float) | ✅ | 492 | CC BY-NC-ND 4.0 | 2025-11-10 | ❌ | Flow matching 说话肖像（非商用） |
| 12 | MuseTalk | [TMElyralab/MuseTalk](https://github.com/TMElyralab/MuseTalk) | ✅ | 6.6k | MIT | 2025-09-26 | ❌ | 唇形同步；**最新 1.5，无 2.0** |
| 13 | LatentSync | [bytedance/LatentSync](https://github.com/bytedance/LatentSync) | ✅ | 6.1k | Apache-2.0 | 2025-06-20 | ❌ | 唇形同步；**最新 1.6，无 2.x** |
| A1 | SoulX-FlashTalk | [Soul-AILab/SoulX-FlashTalk](https://github.com/Soul-AILab/SoulX-FlashTalk) | ✅ | 1.5k | Apache-2.0 | 2026-07-30 | ✅ | **流式 avatar，0.87s 延迟 / 32 FPS** |
| A2 | SoulX-FlashHead | [Soul-AILab/SoulX-FlashHead](https://github.com/Soul-AILab/SoulX-FlashHead) | ✅ | 1.1k | Apache-2.0 | 2026-05 | ✅ | **流式说话头，4090 单卡 96 FPS** |
| A3 | LiveAvatar | [Alibaba-Quark/LiveAvatar](https://github.com/Alibaba-Quark/LiveAvatar) | ✅ | 2.4k | Apache-2.0 | 2026-08-24 | ✅ | **流式无限长，45 FPS（H800）/48GB 卡** |
| A4 | Linly-Talker-Stream | [Kedreamix/Linly-Talker-Stream](https://github.com/Kedreamix/Linly-Talker-Stream) | ✅ | 133 | Apache-2.0 | 2026-02 | ✅ | 全双工流式对话数字人（系统集成） |
| A5 | SentiAvatar | [SentiAvatar/SentiAvatar](https://github.com/SentiAvatar/SentiAvatar) | ✅ | 453 | 未指定 | 2026-04 | ✅ | 交互式 3D 数字人框架 |

---

# 🎯 关键结论

1. **13 个候选名中，12 个能对应到真实公开仓库；只有 "OmniHuman-1.5" 是彻底闭源 API-only**（GitHub 无仓库、HF 无官方权重、仅有 fal.ai / Replicate 端点）。
2. **两个版本号是虚构/未找到的**：
   - **MuseTalk 2.0 → 未找到**（官方最新 1.5，2025-03-28 发布；HF 仅 `TMElyralab/MuseTalk`）
   - **LatentSync 2.x → 未找到**（官方最新 1.6，2025-06-11 发布；HF 仅 `ByteDance/LatentSync` / `-1.5` / `-1.6`）
3. **仓库名易错点**：`Fantasy-AMAP/fantasy-talking2`（全小写）、`jixiaozhong/Sonic`（**不是** tencent-ailab）、`memoavatar/MEMO`、`inclusionAI/Realtime-Venus`、`Alibaba-Quark/LiveAvatar`、`fudan-generative-vision/hallo4`（默认分支 **master**）。
4. **2026 年仍有提交的只有 7 个**：Realtime-Venus(2026-09-21)、Wan2.2(2026-09-21)、Wan-Dancer(2026-07-17)、EchoMimicV3(2026-03-18)、Sonic(2026-01-08)，加上额外的 SoulX-FlashTalk(2026-07-30)、LiveAvatar(2026-08-24)、SoulX-FlashHead(2026-05)、Linly-Talker-Stream(2026-02)、SentiAvatar(2026-04)。
5. **"实时/流式"真正对口的是额外发现的三项**：SoulX-FlashHead（消费级 4090 单卡 96 FPS / 1.3B）、SoulX-FlashTalk（14B / 0.87s / 32 FPS @8×H800）、LiveAvatar（14B / 45 FPS @多卡 H800，48GB 单卡可推理）。原候选清单里的 Ditto（"Realtime"）虽标称实时，但已 **10 个月无更新**。
6. **许可红线**：Sonic 是 **CC BY-NC-SA 4.0**、FLOAT 是 **CC BY-NC-ND 4.0**（均**禁商用**）；HunyuanVideo-Avatar 是**腾讯混元社区许可**（非 SPDX 标准）；Hallo4 / FantasyTalking2 / SentiAvatar **无明确许可证**——商用前必须逐一确认。
7. **Hallo4 性价比极低**：38 star、5 commit、无 LICENSE、HF 权重**门控（401）**、需 H100 级测试环境，且 2025-11 后停更。

---

# 📚 来源清单

**代码仓库（均于 2026-09-22 实测 HTTP 200）**
- https://github.com/inclusionAI/Realtime-Venus
- https://github.com/Wan-Video/Wan-Dancer
- https://github.com/Wan-Video/Wan2.2
- https://github.com/fudan-generative-vision/hallo3
- https://github.com/fudan-generative-vision/hallo4
- https://github.com/antgroup/echomimic_v3
- https://github.com/Tencent-Hunyuan/HunyuanVideo-Avatar
- https://github.com/Fantasy-AMAP/fantasy-talking2
- https://github.com/antgroup/ditto-talkinghead
- https://github.com/memoavatar/MEMO
- https://github.com/jixiaozhong/Sonic
- https://github.com/deepbrainai-research/float
- https://github.com/TMElyralab/MuseTalk
- https://github.com/bytedance/LatentSync
- https://github.com/Soul-AILab/SoulX-FlashTalk
- https://github.com/Soul-AILab/SoulX-FlashHead
- https://github.com/Alibaba-Quark/LiveAvatar
- https://github.com/Kedreamix/Linly-Talker-Stream
- https://github.com/SentiAvatar/SentiAvatar

**不存在（404 实测）**
- https://github.com/Tencent-Hunyuan/Hunyuan-Avatar 类型名（未找到）
- https://github.com/tencent-ailab/Sonic （404）
- https://github.com/bytedance/OmniHuman-1.5 （404，及 OmniHuman / OmniHuman-1 / omnihuman 均 404）
- https://github.com/Fantasy-AMAP/FantasyTalking2 （404，真实名为小写）
- https://github.com/digital-human/ditto-talkinghead （404）
- https://github.com/fudan-generative-vision/hallo4 的 `main` 分支 README （404，默认分支为 master）

**模型权重 / 论文 / 项目页**
- https://huggingface.co/inclusionAI/Realtime-Venus
- https://arxiv.org/abs/2609.13814 · https://realtime-venus.github.io/
- https://huggingface.co/Wan-AI/Wan-Dancer-14B · https://arxiv.org/abs/2607.09581 · https://humanaigc.github.io/wan-dancer-project/
- https://huggingface.co/Wan-AI/Wan2.2-S2V-14B · https://humanaigc.github.io/wan-s2v-webpage
- https://huggingface.co/fudan-generative-ai/hallo4 （门控 401）
- https://dl.acm.org/doi/full/10.1145/3757377.3763914 · https://ar5iv.labs.arxiv.org/html/2505.23525
- https://huggingface.co/BadToBest/EchoMimicV3 · https://huggingface.co/BadToBest/EchoMimicV3/tree/main/echomimicv3-flash-pro · https://arxiv.org/abs/2507.03905 · https://antgroup.github.io/ai/echomimic_v3/
- https://huggingface.co/tencent/HunyuanVideo-Avatar
- https://fal.ai/models/fal-ai/bytedance/omnihuman/v1.5 · https://replicate.com/bytedance/omni-human-1.5 · https://ar5iv.labs.arxiv.org/html/2508.19209
- https://arxiv.org/abs/2508.11255 · https://fantasy-amap.github.io/fantasy-talking2/
- https://huggingface.co/digital-avatar/ditto-talkinghead · https://arxiv.org/abs/2411.19509 · https://digital-avatar.github.io/ai/Ditto/
- https://huggingface.co/memoavatar/memo · https://arxiv.org/abs/2412.04448 · https://memoavatar.github.io
- https://jixiaozhong.github.io/Sonic/ · https://openaccess.thecvf.com/content/CVPR2025/papers/Ji_Sonic_Shifting_Focus_to_Global_Audio_Perception_in_Portrait_Animation_CVPR_2025_paper.pdf
- https://arxiv.org/abs/2412.01064 · https://deepbrainai-research.github.io/float/
- https://huggingface.co/TMElyralab/MuseTalk · https://arxiv.org/abs/2410.10122
- https://huggingface.co/ByteDance/LatentSync-1.6 · https://arxiv.org/abs/2412.09262
- https://huggingface.co/Soul-AILab/SoulX-FlashTalk-14B · https://arxiv.org/pdf/2512.23379
- https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B · https://arxiv.org/pdf/2602.07449 · https://huggingface.co/datasets/Soul-AILab/VividHead
- https://huggingface.co/Quark-Vision/Live-Avatar · https://arxiv.org/abs/2512.04677
- http://finance.people.com.cn/BIG5/n1/2026/0409/c1004-40697921.html · https://m-en.yna.co.kr/view/RPR20260409008000353?section=press-release/index

---

# 📎 附录：父任务侧（同工作区并行核实）补充数据点

以下条目由**父任务侧独立核实**并于 2026-09-22 同步给本报告，**本次未由我二次独立复核**（模型权重体积类数据无法用上面的 shields/Atom 方法验证，故单独列出，标注来源）。

| 项目 | 父任务侧补充数据点 | 我的核实是否一致 |
|---|---|---|
| Realtime-Venus | 只出语音不出脸（9B 全双工对话前端）；**单模型 ~18.7GB bf16** | ✅ 一致（我独立确认仓库/星数/许可/时间；"只出语音"结论一致） |
| Wan-Dancer | **HF Wan-AI/Wan-Dancer-14B 权重 85.67GB**；实测环境 8×A800 80GB | ✅ 环境一致（我从 README 独立读到 8×A800 80GB）；权重体积未独立复核 |
| Wan2.2-S2V | **HF Wan-AI/Wan2.2-S2V-14B，49.15GB** | 体积未独立复核（仓库/权重存在性已独立确认） |
| EchoMimic V3 | AAAI 2026 | ✅ 一致（我另确认 2026-01-22 发布 EchoMimicV3-Flash） |
| HunyuanVideo-Avatar | 自定义许可；**官方最低 24GB** | 许可 ✅ 一致（腾讯混元社区许可）；24GB 门槛未独立复核（README 另载社区 Wan2GP 可单卡 10GB） |
| OmniHuman-1.5 | `bytedance/OmniHuman*` 全 404，仅 Replicate/fal.ai API | ✅ 完全一致 |
| Sonic | last commit **2026-01** | ✅ 一致（Atom feed 精确到 2026-01-08） |
| MuseTalk | 官方 News 2025-03-28；**2.0 未找到** | ✅ 完全一致 |
| LatentSync | 官方 News 2025-06-11，**18GB 门槛**；**2.x 未找到** | ✅ 一致（18GB 我从 README 独立读到；我另确认 HF 仅 `ByteDance/LatentSync`/`-1.5`/`-1.6`） |
| Ditto | ACM MM 2025 | ✅ 一致（我另确认 2025-11-12 开源训练代码） |
| MEMO | last commit 2025-08 | ✅ 一致（Atom feed 精确到 2025-08-06；另确认仅开源 preview 模型） |
| FantasyTalking2 | last commit **2025-08**、2026 无活动 | ✅ 一致（Atom feed 精确到 2025-08-18） |
| Hallo3 / Hallo4 | hallo3 last commit 2025-03；hallo4 仅 38⭐、last commit 2025-11 | ✅ 完全一致（另确认 hallo4 权重门控 401、无 LICENSE） |
| FLOAT | last commit 2025-11 | ✅ 一致（Atom feed 精确到 2025-11-10；另确认 CC BY-NC-ND 4.0） |

**数据来源（badge/API，非人写）**
- https://img.shields.io/github/stars/&lt;owner&gt;/&lt;repo&gt;.json（stars，2026-09-22 实测）
- https://img.shields.io/github/license/&lt;owner&gt;/&lt;repo&gt;.json（license，2026-09-22 实测）
- https://img.shields.io/github/last-commit/&lt;owner&gt;/&lt;repo&gt;.json（月份粒度，2026-09-22 实测）
- https://github.com/&lt;owner&gt;/&lt;repo&gt;/commits/&lt;branch&gt;.atom（**精确提交时间**，2026-09-22 实测）
- https://hf-mirror.com/api/models?search=&lt;kw&gt;（HF 模型版本存在性，2026-09-22 实测）
