_raw_TTS_digitalhuman_integration.md

实时数字人 / 说话头框架的 TTS 原生支持调研(2026-09-22)

调研日期:2026-09-22 调研方法:全部结论来自实际 curl 抓取的原始文件(GitHub raw / jsDelivr / 官方文档站 / 官方配置与源码),逐条引用原文。星标 / 许可证 / 最后提交通过 shields.io JSON 获取(GitHub API 限流,未使用)。 重要说明:

  • 网络过程中 raw.githubusercontent.com 与 github.com 多次超时中断,因此对长文件同时使用了 cdn.jsdelivr.net 与 data.jsdelivr.com(jsDelivr 文件树 API)。凡使用 jsDelivr 的位置均已标注。
  • 凡无法确证的项一律写 未查到,不做推测。
  • 本文件所有 URL 均为本次实际成功抓取的地址。

0. 总览表(先看结论)

框架仓库星标许可证最后提交原生 TTS 引擎
LiveTalkinglipku/LiveTalking9.6kApache-2.0septemberedgetts / gpt-sovits / cosyvoice / fishtts / tencent / doubao / indextts2 / azuretts / qwentts(config.yaml 原注释);代码另有 xtts、omnitts
OpenAvatarChatHumanAIGC-Engineering/OpenAvatarChat3.8kApache-2.0julyEdgeTTS / CosyVoice(本地)/ 百炼 CosyVoice(API);Qwen-Omni 模式下 TTS 由 Qwen-Omni 本体承担
Fayxszyou/Fay14kGPL-3.0yesterdayazure / ali / gptsovits / volcano / gptsovits_v3
Linly-TalkerKedreamix/Linly-Talker3.5kMITfebruaryEdge TTS / PaddleTTS / GPT-SoVITS(推荐)/ XTTS / CosyVoice
Duix.Heygem(现 Duix.Avatar)duixcom/Duix.Heygem → duixcom/Duix.Avatar16knot identifiable by githubaprilfish-speech-ziming(guiji2025/fish-speech-ziming Docker 镜像),固定内置、无切换项
aigcpanelmodstart-lib/aigcpanel5.6kApache-2.0last fridayCosyVoice-300M / CosyVoice-300M-Instruct / CosyVoice2-0.5b / FishSpeech / IndexTTS / SparkTTS / GPT-SoVITS
SoulX-FlashTalkSoul-AILab/SoulX-FlashTalk1.5kApache-2.0july无(纯音频驱动,不含任何 TTS)
Ultralight-Digital-Humananliyuan/Ultralight-Digital-Human2.6knot specifiedjuly无(纯音频特征驱动,不含任何 TTS)
Open-LLM-VTuber(参考)Open-LLM-VTuber/Open-LLM-VTuber14knot identifiable by githubmaysherpa-onnx / pyttsx3 / MeloTTS / Coqui-TTS / GPTSoVITS / Bark / CosyVoice / Edge TTS / Fish Audio / Azure TTS 等
Awesome 类清单weihaox/awesome-digital-human———未查到(该清单聚焦 3D 数字人/avatar 生成,正文未列 TTS 后端)
Awesome 类清单icemaple77/digital-human———edge(免费)/ GPT-SoVITS(自训音色)

1. LiveTalking(lipku/LiveTalking)

1.1 仓库元数据(shields.io)

项值
仓库https://github.com/lipku/LiveTalking
星标{"label":"stars","message":"9.6k",...}
许可证{"label":"license","message":"Apache-2.0",...}
最后提交{"label":"last commit","message":"september",...}

抓取地址:

1.2 README 中的 TTS 描述(逐字引用)

来源:https://raw.githubusercontent.com/lipku/LiveTalking/main/README.md(另抓 README-EN.md,同一提交)

"3. 系统架构 / 逻辑层":

  • TTS 引擎: 模块化设计,支持 EdgeTTS、GPT-SoVITS、CosyVoice、腾讯云等多种方案

英文版同源文件 https://raw.githubusercontent.com/lipku/LiveTalking/main/README-EN.md 第 142 行逐字对应(英文句末还有 "and more"):

  • TTS Engine: Modular design supporting EdgeTTS, GPT-SoVITS, CosyVoice, Tencent Cloud, and more

"插件系统":

  • 基于 registry.py 的去中心化注册机制,开发者可自行扩展 TTS、Avatar、Output 模块

Features:

  1. 支持多种数字人模型: ernerf、musetalk、wav2lip、Ultralight-Digital-Human
  2. 支持声音克隆
  3. 支持数字人说话被打断

核心流程:

核心流程:用户输入文字/音频 → LLM 生成回复(可选)→ TTS 合成语音 → 数字人实时口型同步 → 音视频推流输出

1.3 权威 TTS 列表 = config.yaml(源码级,最可靠)

来源:https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/config.yaml (等价 raw:https://raw.githubusercontent.com/lipku/LiveTalking/main/config.yaml)

# -------- TTS ----------------------------------------------------------
tts: edgetts                   # edgetts / gpt-sovits / cosyvoice / fishtts / tencent / doubao / indextts2 / azuretts / qwentts
REF_FILE: ''                   # voice name or reference audio file path (wav, 16kHz, mono)
REF_TEXT: ''                   # reference text for voice cloning (English or Chinese)
TTS_SERVER: ''                 # TTS server URL, no trailing slash

→ 文档化的官方 TTS 枚举(9 个):edgetts、gpt-sovits、cosyvoice、fishtts、tencent、doubao、indextts2、azuretts、qwentts。

1.4 代码级 TTS 模块清单(比 config.yaml 注释多出 2 个)

文件树来源:https://data.jsdelivr.com/v1/packages/gh/lipku/LiveTalking@main

/tts/__init__.py
/tts/azure.py
/tts/base_tts.py
/tts/cosyvoice.py
/tts/doubao.py
/tts/edge.py
/tts/fish.py
/tts/indextts2.py
/tts/omnitts.py
/tts/qwentts.py
/tts/sovits.py
/tts/tencent.py
/tts/xtts.py

已从源码中逐一核实的注册名(@register("tts", "<name>") 装饰器):

文件注册名抓取地址(jsDelivr CDN)
tts/edge.pyedgettshttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/edge.py
tts/azure.pyazuretts.../tts/azure.py
tts/doubao.pydoubao.../tts/doubao.py
tts/omnitts.pyomnitts.../tts/omnitts.py
tts/qwentts.pyqwentts.../tts/qwentts.py
tts/sovits.pygpt-sovits.../tts/sovits.py
tts/tencent.pytencent.../tts/tencent.py
tts/xtts.pyxtts.../tts/xtts.py
tts/cosyvoice.py未取到(CDN 404,但文件树中存在).../tts/cosyvoice.py
tts/fish.py未取到(同上).../tts/fish.py
tts/indextts2.py未取到(同上).../tts/indextts2.py

说明:cosyvoice.py / fish.py / indextts2.py 三个文件在 jsDelivr 文件树 API 中存在,但 CDN 直取返回 404(jsDelivr 的 tree 与 CDN 可能解析到不同提交),因此其注册名未能直接引用原文;对应的 config 名(cosyvoice / fishtts / indextts2)来自上面 config.yaml 的原文注释。 另注:xtts 与 omnitts 存在于代码,但未出现在 config.yaml 的注释枚举里。

1.5 是否可插拔?——是(去中心化注册表)

来源:https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/registry.py

_REGISTRY: Dict[str, Dict[str, Type]] = {
    "stt": {}, "llm": {}, "tts": {}, "avatar": {}, "output": {},
}

def register(category: str, name: str):
    """
    装饰器:注册插件类到全局注册表。

    用法::

        @register("tts", "edgetts")
        class EdgeTTS(BaseTTS): ...
    """

以及 create(category, name, **kwargs) 的报错文案原文:

f"Plugin '{name}' not found in category '{category}'. Available: {available}"

→ 结论:TTS 层完全可插拔。任意自定义 TTS 只需继承 BaseTTS 并加 @register("tts", "<name>"),再用 tts: <name> 启用。

1.6 延迟 / 流式要求

1.7 输出格式要求(明确)

来源:https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/base_tts.py

class BaseTTS:
    def __init__(self, opt, parent: "BaseAvatar"):
        self.opt = opt
        self.parent = parent

        #self.fps = opt.fps # 20 ms per frame
        self.sample_rate = 16000
        self.chunk = self.sample_rate // (opt.fps*2) # 320 samples per chunk (20ms * 16000 / 1000)

1.8 官方 TTS 部署文档(doc.livetalking.ai)

来源:https://doc.livetalking.ai/docs/tts/

页面为 TTS 索引,正文逐字:

TTS 相关文档索引。 GPT-SoVITS Fish Speech CosyVoice 旧版说明

已抓取的子页:

子页URL要点(逐字)
GPT-SoVITShttps://doc.livetalking.ai/docs/tts/gptsovits/"采用gpt-sovits方案,bert-sovits适合长音频训练,gpt-sovits运行短音频快速推理";endpoint: /tts;示例含 streaming_mode=true;"media_type" : "wav" # support "wav" , "raw" , "ogg" , "aac"
Fish Speechhttps://doc.livetalking.ai/docs/tts/fishspeech/endpoint: /v1/tts;"chunk_length" : 200、"format" : "wav"、"streaming" : false(默认)、fishaudio/fish-speech-1.5
CosyVoicehttps://doc.livetalking.ai/docs/tts/cosyvoice/用 iic/CosyVoice2-0.5B + iic/CosyVoice-ttsfrd;启动 CosyVoice/runtime/python/fastapi/server.py --model_dir .../CosyVoice2-0.5B
OmniTTShttps://doc.livetalking.ai/docs/tts/omnitts/一个 OpenAI 兼容的统一 TTS 服务(vLLM-Omni),覆盖多个模型——见下

OmniTTS 页逐字要点(这是 LiveTalking 官方"一站式接多 TTS"的答案):

vLLM-Omni是一个推理服务系统,已经支持了多种tts模型,包括qwen3-tts、cosyvoice、VoxCpm2。更多模型 https://docs.vllm.ai/projects/vllm-omni/en/latest/serving/speech_api/ 。对外提供统一的tts API接口服务,兼容openai规范,支持克隆声音的上传、删除、查看。

其 "3. 启动api服务" 小节给出的模型与显存/延迟(逐字摘录):

小节命令原文备注
3.1 Qwen3vllm serve Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice --omni --trust-remote-code --port 8091"需要显存12g"
3.2 Voxcpmvllm serve openbmb/VoxCPM2 --omni --trust-remote-code --port 8091"需要显存13G"
3.3 Fish Speechvllm serve fishaudio/s2-pro --omni --trust-remote-code --port 8091"需要显存18G"
3.4 IndexTTSvllm serve IndexTeam/IndexTTS-2 --omni --trust-remote-code --port 8091"需要显存17G,延时4s"
3.5 CosyVoice(uv pip install s3tokenizer …)"需要显存14G"

1.9 FAQ / 其他

https://doc.livetalking.ai/docs/faq/ 抓取成功,但正文仅见 pytorch3d 安装问题,未涉及 TTS 延迟或格式 → TTS 相关 Q&A 未查到。


2. OpenAvatarChat(HumanAIGC-Engineering/OpenAvatarChat)

⚠️ 重要更正:任务给出的 HumanAIGC/OpenAvatarChat 不存在(shields.io 返回 "message":"repo not found")。实际仓库所有者为 HumanAIGC-Engineering:https://github.com/HumanAIGC-Engineering/OpenAvatarChat。 抓取 https://raw.githubusercontent.com/HumanAIGC/OpenAvatarChat/main/README.md 亦失败(非 200)。

2.1 仓库元数据(shields.io)

项值
仓库https://github.com/HumanAIGC-Engineering/OpenAvatarChat
星标{"label":"stars","message":"3.8k",...}
许可证{"label":"license","message":"Apache-2.0",...}
最后提交{"label":"last commit","message":"july",...}

抓取地址:https://img.shields.io/github/{stars,license,last-commit}/HumanAIGC-Engineering/OpenAvatarChat.json

2.2 README 中的 TTS 描述(逐字引用)

来源:https://raw.githubusercontent.com/HumanAIGC-Engineering/OpenAvatarChat/main/README.md

核心亮点:

  • 模块化架构设计:采用高度模块化设计,可灵活替换 ASR、LLM、TTS、Avatar 等核心组件
  • 低延迟优化:通过 VAD 检测、语音缓冲、帧率控制等机制优化,平均响应时间仅 2.2 秒

组件依赖表(TTS 行):

| TTS | FunAudioLLM/CosyVoice |[GitHub]](https://github.com/FunAudioLLM/CosyVoice)||

预置模式表(TTS 列):

| CONFIG名称 | ASR | LLM | TTS | AVATAR | | chat_with_lam.yaml | SenseVoice | API | API | LAM | | chat_with_qwen_omni.yaml | Qwen-Omni | Qwen-Omni | Qwen-Omni | lite-avatar | | chat_with_openai_compatible_bailian_cosyvoice.yaml | SenseVoice | API | API | lite-avatar | | chat_with_openai_compatible_bailian_cosyvoice_flashhead.yaml | SenseVoice | API | API | FlashHead |

最新动态(2026.04,v0.6.0):

  • [2026.04] ⭐️⭐️⭐️ 版本 0.6.0发布: … 接入 SoulX-FlashHead 数字人,基于扩散模型的实时流式说话头生成

2.3 官方文档的 TTS Handler 清单(权威)

来源:https://humanaigc-engineering.github.io/OpenAvatarChat/reference/preset-modes

侧边栏 "TTS(语音合成)" 下只有三项,逐字:

TTS(语音合成) 百炼 CosyVoice CosyVoice 本地 Edge TTS

对应文档页(均抓取成功):

Handler文档 URL关键点(逐字)
百炼 CosyVoicehttps://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/bailian-cosyvoice.htmlmodel_name : "cosyvoice-v1";TTS_CosyVoice.sample_rate | 24000 | 输出音频采样率
CosyVoice 本地https://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/cosyvoice-local.html"TTS 默认为 CosyVoice 的 iic/CosyVoice-300M-SFT + 中文女,可以通过修改为其他模型配合 ref_audio_path 和 ref_audio_text 进行音色复刻。"
Edge TTShttps://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/edge-tts.html"集成微软的 Edge TTS,使用云端推理,无需申请 API Key。" 配置:Edge_TTS: module: tts/edgetts/tts_handler_edgetts / voice: "zh-CN-XiaoxiaoNeural"

2.4 代码级 TTS Handler 清单(与文档一致)

文件树来源:https://data.jsdelivr.com/v1/packages/gh/HumanAIGC-Engineering/OpenAvatarChat@main

/src/handlers/tts/bailian_tts/tts_handler_cosyvoice_bailian.py
/src/handlers/tts/cosyvoice/tts_handler_cosyvoice.py
/src/handlers/tts/cosyvoice/cosyvoice_processor.py
/src/handlers/tts/edgetts/tts_handler_edgetts.py
/config/chat_with_openai_compatible_edge_tts.yaml
/config/chat_with_openai_compatible_bailian_cosyvoice.yaml
/config/chat_with_gs.yaml

→ 确认 OpenAvatarChat 原生 TTS 只有 3 条路径:本地 CosyVoice、百炼 CosyVoice API、Edge TTS。(chat_with_qwen_omni.yaml 模式下 TTS 由 Qwen-Omni 多模态模型本体输出,不算独立 TTS 引擎。)

2.5 是否可插拔?——是(Handler 插件架构)

来源:https://cdn.jsdelivr.net/gh/HumanAIGC-Engineering/OpenAvatarChat@main/config/chat_with_openai_compatible_edge_tts.yaml

  chat_engine:
    model_root: "models"
    concurrent_limit: 1
    handler_search_path:
      - "src/handlers"
    handler_configs:
      Edge_TTS:
        enabled: True
        module: tts/edgetts/tts_handler_edgetts
        voice: "zh-CN-XiaoxiaoNeural"

→ 通过 handler_search_path + module: 声明式装配,新增 TTS 只要按 HandlerBase 写一个模块即可。(框架无"任意 TTS 免代码接入"的网关,需要写 handler。)

2.6 延迟 / 流式要求

2.7 输出格式要求

来源:https://cdn.jsdelivr.net/gh/HumanAIGC-Engineering/OpenAvatarChat@main/src/handlers/tts/cosyvoice/tts_handler_cosyvoice.py 与 .../edgetts/tts_handler_edgetts.py

两个 Handler 的 TTSConfig 都声明:

sample_rate: int = Field(default=24000)
sample_rate: int = Field(default=24000)

输出为单声道(DataBundleEntry.create_audio_entry("avatar_audio", 1, self.sample_rate) 的第 2 个参数 1 = 通道数):

definition.add_entry(DataBundleEntry.create_audio_entry("avatar_audio", 1, self.sample_rate))

3. Fay(xszyou/Fay)

3.1 仓库元数据(shields.io)

项值
仓库https://github.com/xszyou/Fay
星标{"label":"stars","message":"14k",...}
许可证{"label":"license","message":"GPL-3.0",...}
最后提交{"label":"last commit","message":"yesterday",...}

抓取地址:https://img.shields.io/github/{stars,license,last-commit}/xszyou/Fay.json

3.2 README 中的 TTS 描述(逐字引用)

来源:https://raw.githubusercontent.com/xszyou/Fay/master/README.md(注意默认分支为 master)

Fay数字人框架,向上适配各种数字人模型技术,向下接入各式大语言模型,并且便于更换诸如TTS、ASR等模型,为单片机、app、网站提供全面的数字人应用接口。

功能特点:

  • 自由匹配数字人模型、大语言模型(openai 兼容接口)、ASR、TTS模型
  • 全时流式的支持
  • 支持唤醒及打断对话

⚠️ README 正文没有列出具体 TTS 引擎名称,只给飞书文档链接。因此下面的引擎清单来自源码配置(更权威)。

3.3 权威 TTS 引擎清单 = system.conf.bak(逐字引用)

来源:https://raw.githubusercontent.com/xszyou/Fay/master/system.conf.bak

#tts类型(切换请重新选择所需要的声音)azure、ali、gptsovits、volcano、gptsovits_v3
tts_module=ali

# 微软 文字转语音 服务密钥(非必须,使用可产生不同情绪的音频)https://azure.microsoft.com/zh-cn/services/cognitive-services/text-to-speech/
ms_tts_key=
ms_tts_region=

# 阿里云 文字转语音 服务密钥 https://ai.aliyun.com/nls/trans
ali_tss_key_id=
ali_tss_key_secret=
ali_tss_app_key=

# Doubao-语音合成 服务密钥 https://www.volcengine.com/product/voice-tech
volcano_tts_appid=
volcano_tts_access_token=
volcano_tts_cluster=volcano_tts
#可为空,为空时读取选择的音色
volcano_tts_voice_type=

→ Fay 官方 TTS 引擎(5 个):azure(微软)、ali(阿里云 NLS)、gptsovits、volcano(火山引擎/豆包语音)、gptsovits_v3。

旁证(DeepWiki,索引时间 "Last indexed: 29 March 2026"):https://deepwiki.com/xszyou/Fay/10.1-local-setup-and-configuration

| tts_module | Text-to-speech engine | ali , azure , volcano , gptsovits |

(DeepWiki 该表未列 gptsovits_v3;system.conf.bak 为更权威来源。)

3.4 TTS 相关源码文件

来源:https://raw.githubusercontent.com/xszyou/Fay/master/tts/ali_tss.py、.../tts/ms_tts_sdk.py、.../tts/tts_voice.py (三者与 jsDelivr 版本 https://cdn.jsdelivr.net/gh/xszyou/Fay@master/tts/* 内容一致,已逐行 diff 一致)

import azure.cognitiveservices.speech as speechsdk
import edge_tts
...
if config_util.key_ms_tts_key and config_util.key_ms_tts_key is not None and config_util.key_ms_tts_key.strip() != "":
    self.__speech_config = speechsdk.SpeechConfig(subscription=cfg.key_ms_tts_key, region=cfg.key_ms_tts_region)
    ...
    self.__speech_config.set_speech_synthesis_output_format(speechsdk.SpeechSynthesisOutputFormat.Riff16Khz16BitMonoPcm)
    self.__synthesizer = speechsdk.SpeechSynthesizer(speech_config=self.__speech_config, audio_config=None)
    self.ms_tts = True
class EnumVoice(Enum):
    XIAO_XIAO_NEW = {
        "name": "晓晓(azure)",
        "voiceName": "zh-CN-XiaoxiaoMultilingualNeural",
        "styleList": {"angry": "angry", "lyrical": "lyrical", "calm": "gentle",
                      "assistant": "affectionate", "cheerful": "cheerful"}
    }
    XIAO_XIAO = { "name": "晓晓(edge)", "voiceName": "zh-CN-XiaoxiaoNeural", ... }

DeepWiki TTS 页(https://deepwiki.com/xszyou/Fay/4.2-text-to-speech-(tts))补充(逐字):

Azure SDK : If key_ms_tts_key is provided, it uses the official speechsdk for high-fidelity synthesis with SSML support for fine-grained style control Edge TTS (Fallback) : If no API key is found, it falls back to the edge_tts library, which provides free access to Microsoft's neural voices via the Edge browser's API Format Conversion : Edge TTS outputs .mp3 , which is automatically converted to .wav using pydub to ensure compatibility with the Fay audio player Audio Processing : Converts the returned stream into a WAV file with 16kHz sampling rate and mono channel

3.5 是否可插拔?——是(配置切换),但不是任意 TTS 免代码

3.6 延迟 / 流式要求

3.7 输出格式要求


4. Linly-Talker(Kedreamix/Linly-Talker)

⚠️ 重要更正:任务给出的 Korvo-AI/Linly-Talker 不存在(shields.io 返回 "message":"repo not found")。实际仓库为 Kedreamix/Linly-Talker。

4.1 仓库元数据(shields.io)

项值
仓库https://github.com/Kedreamix/Linly-Talker
星标{"label":"stars","message":"3.5k",...}
许可证{"label":"license","message":"MIT",...}
最后提交{"label":"last commit","message":"february",...}

抓取地址:https://img.shields.io/github/{stars,license,last-commit}/Kedreamix/Linly-Talker.json

4.2 README 的 TTS 目录(逐字引用)

来源:https://raw.githubusercontent.com/Kedreamix/Linly-Talker/main/README.md

README 目录树原文(第 87–99 行):

  - [ASR - Speech Recognition](#asr---speech-recognition)
    - [Whisper](#whisper)
    - [FunASR](#funasr)
    - [Coming Soon](#coming-soon)
  - [TTS - Text To Speech](#tts---text-to-speech)
    - [Edge TTS](#edge-tts)
    - [PaddleTTS](#paddletts)
    - [Coming Soon](#coming-soon-1)
  - [Voice Clone](#voice-clone)
    - [GPT-SoVITS(Recommend)](#gpt-sovitsrecommend)
    - [XTTS](#xtts)
    - [CosyVoice](#cosyvoice)
    - [Coming Soon](#coming-soon-2)

→ 文档化的 TTS / 声音克隆后端(5 个):Edge TTS、PaddleTTS、GPT-SoVITS(标注 Recommend)、XTTS、CosyVoice。 (README 把前两者归在 "TTS",后三者归在 "Voice Clone",但四者都是语音合成实现。)

关键原文摘录:

Edge TTS

To use Microsoft Edge's online text-to-speech service from Python without needing Microsoft Edge or Windows or an API key, you can refer to ... https://github.com/rany2/edge-tts

Due to some issues with the Edge TTS repository, it seems that Microsoft has restricted certain IPs. ... I recommend using the CosyVoice method.

PaddleTTS

In practical use, there may be scenarios that require offline operation. Since Edge TTS requires an online environment to generate speech, we have chosen PaddleSpeech, another open-source alternative, for Text-to-Speech (TTS). ... https://github.com/PaddlePaddle/PaddleSpeech

GPT-SoVITS(Recommend)

Thank you for your open source contribution. I have also found the GPT-SoVITS voice cloning model to be quite impressive. ... https://github.com/RVC-Boss/GPT-SoVITS

XTTS

Coqui XTTS is a leading deep learning toolkit for Text-to-Speech (TTS) tasks, allowing for voice cloning and voice transfer to different languages using a 5-second or longer audio clip.

CosyVoice

CosyVoice is an open-source multilingual speech understanding model developed by Alibaba's Tongyi Lab ... 1. CosyVoice-300M ... 2. CosyVoice-300M-SFT ... 3. CosyVoice-300M-Instruct ...

更新日志中的相关条目:

  • **Updated the offline mode for Paddle TTS, excluding Edge TTS.**
  • **Implemented a simple fix for the Edge-TTS bug, resolved several issues with MuseTalk, and plan to integrate fishTTS for more stable TTS performance**

("plan to integrate fishTTS" —— 计划中,非已完成。)

架构定位(原文):

  • 🧩 Modular Multimodal Pipeline: Reuses existing ASR/LLM/TTS/Avatar capabilities while adopting streaming processing framework for system refactoring.

4.3 子文档 TTS/README.md

来源:https://cdn.jsdelivr.net/gh/Kedreamix/Linly-Talker@main/TTS/README.md(3433 B,同时以 raw 路径核对过)

标题逐字:

TTS 赋予数字人真实的语音交互能力

Edge-TTS

PaddleTTS

正文包含自实现的 EdgeTTS 包装类(SUPPORTED_VOICE、predict(...) 输出 result.wav + result.vtt 字幕)与 PaddleSpeech 的 TTSExecutor 封装(am='fastspeech2'/'tacotron2'、voc='pwgan'/'hifigan'/...、lang='zh'/'en'/'mix'/'canton')。

4.4 是否可插拔?——部分,需改代码

4.5 延迟 / 流式要求

4.6 输出格式要求


5. Duix.Heygem(duixcom/Duix.Heygem → 现名 duixcom/Duix.Avatar)

5.1 仓库元数据(shields.io)

项Duix.HeygemDuix.Avatar
星标{"label":"stars","message":"16k",...}{"label":"stars","message":"16k",...}
许可证{"label":"license","message":"not identifiable by github",...}同左
最后提交{"label":"last commit","message":"april",...}{"label":"last commit","message":"april",...}

抓取地址:https://img.shields.io/github/{stars,license,last-commit}/duixcom/Duix.Heygem.json 与 .../duixcom/Duix.Avatar.json

→ 两仓库三项指标完全一致,且 README 内容即 "Duix.Avatar";README 内 LICENSE 链接指向 https://github.com/duixcom/Duix.Avatar/blob/main/LICENSE。结论:Duix.Heygem 已改名为 Duix.Avatar(同一仓库)。

5.2 README 中的 TTS(逐字引用)

来源:https://raw.githubusercontent.com/duixcom/Duix.Heygem/main/README.md

依赖(Docker 镜像):

  1. Nodejs 18
  2. Docker Images
    • docker pull guiji2025/fun-asr
    • docker pull guiji2025/fish-speech-ziming

致谢章节(最直接的一句):

10. Acknowledgments

  • ASR based on fun-asr
  • TTS based on fish-speech-ziming

关键路径:

  • src/main/service/voice.js
  1. Separate video into silent video + audio
  2. Place audio in D:\duix_avatar_data\voice\data is agreed with the guiji2025/fish-speech-ziming service, can be modified in docker-compose

→ 原生 TTS 只有 1 个:fish-speech-ziming(GuiJi/硅基智能定制的 Fish-Speech Docker 服务,guiji2025/fish-speech-ziming)。 无其他内置 TTS。

5.3 是否可插拔?——否(未见文档化的 TTS 替换机制)

5.4 延迟 / 流式要求

### **Audio Synthesis**
Interface: `http://127.0.0.1:18180/v1/invoke`

{
  "speaker": "{uuid}", // A unique UUID
  "text": "xxxxxxxxxx", // Text content to synthesize
  "format": "wav", // Fixed parameter
  "topP": 0.7, // Fixed parameter
  "max_new_tokens": 1024, // Fixed parameter
  "chunk_length": 100, // Fixed parameter
  "repetition_penalty": 1.2, // Fixed parameter
  "temperature": 0.7, // Fixed parameter
  "need_asr": false, // Fixed parameter
  "streaming": false, // Fixed parameter
  "is_fixed_seed": 0, // Fixed parameter
  "is_norm": 0, // Fixed parameter
  "reference_audio": "{voice.asr_format_audio_url}", // Return value from previous "Model Training" step
  "reference_text": "{voice.reference_audio_text}" // Return value from previous "Model Training" step
}

5.5 输出格式要求


6. aigcpanel(modstart-lib/aigcpanel)

owner 确认:正确仓库为 modstart-lib/aigcpanel(抓取成功);README 亦引用 https://aigcpanel.com 与 Gitee 同路径。

6.1 仓库元数据(shields.io)

项值
仓库https://github.com/modstart-lib/aigcpanel
星标{"label":"stars","message":"5.6k",...}
许可证{"label":"license","message":"Apache-2.0",...}
最后提交{"label":"last commit","message":"last friday",...}

抓取地址:https://img.shields.io/github/{stars,license,last-commit}/modstart-lib/aigcpanel.json

6.2 README 中的支持模型(逐字引用)

来源:https://raw.githubusercontent.com/modstart-lib/aigcpanel/main/README.md

软件介绍:

AIGCPanel 是一款简单易用的一站式 AI 数字人桌面应用,支持 Windows / macOS / Linux 三平台。… 核心能力涵盖:数字人视频合成(换口型)、语音合成 / 克隆 / 识别、25+ 音视频处理工具、智能直播互动,Pro 版额外提供可视化工作流编排和云端 AI 模型服务。 软件内置模型市场,支持一键下载启动包,开箱即用;同时兼容远程 API 模型,灵活适配各类部署场景。

"支持模型 → 声音合成"表格(逐字):

声音合成

| 模型 | 说明 | | CosyVoice-300M | 阿里通义实验室开源 TTS | | CosyVoice-300M-Instruct | 指令控制版 | | CosyVoice2-0.5b | 第二代轻量版 | | FishSpeech | 高质量零样本语音克隆 | | IndexTTS | 工业级中文 TTS | | SparkTTS | 讯飞开源语音合成 | | GPT-SoVITS | 少样本声音克隆 |

→ aigcpanel 原生 TTS(7 项):CosyVoice-300M、CosyVoice-300M-Instruct、CosyVoice2-0.5b、FishSpeech、IndexTTS、SparkTTS、GPT-SoVITS。

"声音识别"表格:FunASR(阿里达摩院开源 ASR,支持带时间戳输出) "视频模型"表格:MuseTalk、LatentSync、Wav2Lip、Heygem

6.3 是否可插拔?——是(模型市场 + 远程 API)

README 逐字:

  • 本地模型一键导入、启动/停止、日志查看、参数配置
  • 支持远程 API 模型接入
  • 云端 AI 模型服务(无需本地显卡)(VIP)

代码侧(文件树来源 https://data.jsdelivr.com/v1/packages/gh/modstart-lib/aigcpanel@master)可见 TTS 走统一任务/工作流抽象:

/src/pages/Apps/LongTextTts/...        (长文本转音频)
/src/pages/Apps/SubtitleTts/...        (字幕转音频)
/src/pages/Sound/components/SoundGenerateForm.vue
/src/pages/Sound/components/SoundGenerateSelector.vue   ← 模型选择器
/src/task/SoundGenerate.ts

→ 模型以"启动包/模型市场"形式可插拔,并支持远程 API 模型;但"接入任意第三方 TTS"的自定义插件文档:未查到。

6.4 延迟 / 流式要求

6.5 输出格式要求


7. SoulX-FlashTalk(Soul-AILab/SoulX-FlashTalk)

7.1 仓库元数据(shields.io)

项值
仓库https://github.com/Soul-AILab/SoulX-FlashTalk
星标{"label":"stars","message":"1.5k",...}
许可证{"label":"license","message":"Apache-2.0",...}
最后提交{"label":"last commit","message":"july",...}

抓取地址:https://img.shields.io/github/{stars,license,last-commit}/Soul-AILab/SoulX-FlashTalk.json

7.2 结论:不含任何 TTS 引擎

来源:https://raw.githubusercontent.com/Soul-AILab/SoulX-FlashTalk/main/README.md(12367 B)

全文标题逐字:

SoulX-FlashTalk: Real-Time Infinite Streaming of Audio-Driven Avatars via Self-Correcting Bidirectional Distillation

| SoulX-FlashTalk-14B | Our 14b model | 🤗 Huggingface | | chinese-wav2vec2-base | chinese-wav2vec2-base | 🤗 Huggingface |

7.3 是否可插拔?

7.4 延迟 / 流式要求

Real-time inference speed can only be supported on 8xH800 or higher graphics cards

Requires more than 64G of VRAM. Use --cpu_offload to reduce VRAM usage to 40G.

help="stream: encode audio chunk before every generation; once: encode audio together"

7.5 输入音频格式要求(明确,可反推对 TTS 的要求)

来源:https://cdn.jsdelivr.net/gh/Soul-AILab/SoulX-FlashTalk@main/flash_talk/configs/infer_params.yaml(逐字全文)

frame_num: 33
motion_frames_num: 5
tgt_fps: 25
sample_rate: 16000
sample_steps: 4
sample_shift: 5
color_correction_strength: 1.0
cached_audio_duration: 8
height: 720
width: 416

来源:https://cdn.jsdelivr.net/gh/Soul-AILab/SoulX-FlashTalk@main/generate_video.py(逐字)

human_speech_array_all, _ = librosa.load(args.audio_path, sr=infer_params['sample_rate'], mono=True)
human_speech_array_slice_len = slice_len * sample_rate // tgt_fps
human_speech_array_frame_num = frame_num * sample_rate // tgt_fps
default="examples/cantonese_16k.wav",
help="[meta file] The audio path to generate the video.")

8. Ultralight-Digital-Human(anliyuan/Ultralight-Digital-Human)

8.1 仓库元数据(shields.io)

项值
仓库https://github.com/anliyuan/Ultralight-Digital-Human
星标{"label":"stars","message":"2.6k",...}
许可证{"label":"license","message":"not specified",...}
最后提交{"label":"last commit","message":"july",...}

抓取地址:https://img.shields.io/github/{stars,license,last-commit}/anliyuan/Ultralight-Digital-Human.json

注:raw.githubusercontent.com/.../main/README.md 返回 404;默认分支为 master,取自 https://cdn.jsdelivr.net/gh/anliyuan/Ultralight-Digital-Human@master/README.md。

8.2 结论:不含任何 TTS 引擎

来源:https://cdn.jsdelivr.net/gh/anliyuan/Ultralight-Digital-Human@master/README.md

inference

Before run inference, you need to extract test audio feature(i will merge this step and inference step), run this 在推理之前,需要先提取测试音频的特征(之后会把这步和推理合并到一起去),运行(音频采样率需要是16000)

python data_utils/hubert.py --wav your_test_audio.wav  # when using hubert
or
python data_utils/python wenet_infer.py your_test_audio.wav  # when using wenet

then you get your_test_audio_hu.npy or your_test_audio_wenet.npy

python inference.py --asr hubert --dataset ./your_data_dir/ --audio_feat your_test_audio_hu.npy --save_path xxx.mp4 --checkpoint your_trained_ckpt.pth

To merge the audio and the video, run

ffmpeg -i xxx.mp4 -i your_audio.wav -c:v libx264 -c:a aac result_test.mp4

上游 TTS 完全由用户自备(音频 → 特征 → 视频,再 ffmpeg 合音轨)。

8.3 是否可插拔?

8.4 延迟 / 流式要求

关于流式推理:

使用流式推理时,建议把静音的图片和对应的关键点放在单独的目录里,img_inference和lms_inference里。

因为一般用到流式推理的场景一般对实时性要求比较高,所以这里我只写了wenet作为音频编码器的情况(实测在2080这样的机器上多个并发时每帧音频处理+视频处理耗时10ms以内,需要将模型转为onnx)。

这个模型是支持流式推理的,但是代码还没有完善,之后我会提上来。

c++流式推理代码已开源到FeatherTalk🎉 … 它是 Ultralight Digital Human 的整理和升级版,重点优化了音频编码器、训练流程和移动端部署体验。

8.5 输入音频格式要求(明确)

When you using wenet, you neet to ensure that your video frame rate is 20, and for hubert, your video frame rate should be 25. 如果你选择使用wenet的话,你必须保证你视频的帧率是20fps,如果选择hubert,视频帧率必须是25fps。


9. Open-LLM-VTuber(参考:TTS 后端最丰富的类数字人项目)

9.1 仓库元数据(shields.io)

项值
仓库https://github.com/Open-LLM-VTuber/Open-LLM-VTuber
星标{"label":"stars","message":"14k",...}
许可证{"label":"license","message":"not identifiable by github",...}(README 徽章指向 LICENSE,第三方页脚称 "the MIT license of this project")
最后提交{"label":"last commit","message":"may",...}

抓取地址:https://img.shields.io/github/{stars,license,last-commit}/Open-LLM-VTuber/Open-LLM-VTuber.json

9.2 TTS 清单(逐字引用)

来源:https://raw.githubusercontent.com/Open-LLM-VTuber/Open-LLM-VTuber/main/README.md(11536 B)

🧠 Extensive model support:

  • 🤖 Large Language Models (LLM): Ollama, OpenAI (and any OpenAI-compatible API), Gemini, Claude, Mistral, DeepSeek, Zhipu AI, GGUF, LM Studio, vLLM, etc.
  • 🎙️ Automatic Speech Recognition (ASR): sherpa-onnx, FunASR, Faster-Whisper, Whisper.cpp, Whisper, Groq Whisper, Azure ASR, etc.
  • 🔊 Text-to-Speech (TTS): sherpa-onnx, pyttsx3, MeloTTS, Coqui-TTS, GPTSoVITS, Bark, CosyVoice, Edge TTS, Fish Audio, Azure TTS, etc.

→ 10 个已点名 TTS("etc." 表示还有更多):sherpa-onnx、pyttsx3、MeloTTS、Coqui-TTS、GPTSoVITS、Bark、CosyVoice、Edge TTS、Fish Audio、Azure TTS。

9.3 可插拔性(逐字)

  • 🔧 Highly customizable:
    • ⚙️ Simple module configuration: Switch various functional modules through simple configuration file modifications, without delving into the code
    • 🔌 Good extensibility: Modular design allows you to easily add your own LLM, ASR, TTS, and other module implementations, extending new features at any time

→ 配置文件切换 + 可自行新增 TTS 模块实现(本报告中可插拔性最好的一个)。

9.4 延迟 / 流式 / 格式


10. "Awesome-Digital-Human" 类清单

10.1 weihaox/awesome-digital-human

10.2 icemaple77/digital-human

四个模块(STT / LLM / TTS / 数字人)各自独立,换实现只改 .env 一处,主程序不动。 ├─ TTS 文字→声音 (edge 免费 | GPT-SoVITS 自训音色) | 嘴 TTS | edge(免费无模型)/ gptsovits(自训音色) | TTS_PROVIDER;gptsovits 配 TTS_BASE_URL+参考音频 |

→ TTS 后端:edge(免费、无模型)/ gptsovits(自训音色,走 :9880 api_v2)。 通过 .env 的 TTS_PROVIDER 切换 —— 配置级可插拔(但仅两个内置实现)。


11. 2026 年社区 / 中文博客推荐:数字人该配什么 TTS

11.1 LiveTalking 专用 TTS 选型表(CSDN,2026-01-24 发布 / 2026-03-30 更新)

来源:https://blog.csdn.net/jacke121/article/details/157323875(标题《LiveTalking 部署笔记》)

该文在 LiveTalking 语境下给出「tts选型」表,逐字摘录(含首包延迟与显存):

算法名称核心特点克隆速度与效果资源/硬件要求一句话总结
IndexTTS2效果与资源的完美平衡,支持时长与情感精细控制克隆效果好,可控制情感强度,首包延迟约1.2秒8GB显存 即可流畅运行✅ 非常适合,资源友好,效果出色
FishTTS / FishSpeech速度之王,流式架构带来毫秒级响应首帧延迟低于 500ms,10秒样本即可克隆显存占用仅 3.2GB,轻量高效✅ 非常适合,极速体验,资源占用极低
CosyVoice全能选手,功能全面均衡,由阿里开源克隆相似度达95%以上,流式首包延迟低于 150ms显存需求约 6GB⭐ 可考虑,表现均衡,但非克隆特长生
XTTS多语言专家,一次训练支持17种语言支持多语言零样本克隆,但社区实测克隆相似度略低于专注中文的模型约 6-8GB⚠️ 除非你需要极强的多语言能力,否则不是首选
SoVITS歌声转换,专为"AI翻唱"设计,非普通语音克隆效果惊艳,但需长音频(10-30分钟)训练,推理显存占用大显存占用大,6GB显存仅能处理30秒音频❌ 不适合,场景是唱歌,且资源消耗大
EdgeTTS便捷免费的云端服务,非本地克隆算法无法克隆声音,只能使用微软预设的200多种音色无需显卡,只需网络❌ 不适合,不支持声音克隆

→ 与 LiveTalking 内置列表对照:IndexTTS2 / FishTTS / CosyVoice / XTTS / SoVITS(即 gpt-sovits) / EdgeTTS 全部都在 LiveTalking 的原生枚举里(见 §1.3)。也就是说社区推荐顺序 = IndexTTS2 ≈ FishSpeech > CosyVoice > XTTS > SoVITS/EdgeTTS。 该文另有一条 LiveTalking 实操原文(逐字):

默认的是edgetts,我选择采用gpt-sovits,edgetts相较于gpt-sovits生成的速度会快一点,但不过音色质量会相对弱一点,如果电脑配置支持,可以在本地部署一个gpt-sovits。

以及:

--REF_FILE:用于声音克隆的样例音频文件,--REF_TEXT:样例文本信息,--push_url:srs推流服务地址。

11.2 2026 开源 TTS 横向对比(掘金,2026-06-12,作者 武子康)

来源:https://juejin.cn/post/7649764223339495475 标题逐字:《调查研究-169 开源 TTS 模型横向对比:从"能发声"到"可部署的语音智能基础设施"(2026 版)》

摘要逐字:

场景:面向语音助手、陪伴机器人、短视频配音、有声书、数字人、客服播报等场景的 TTS 选型与工程落地。 结论:当前 TTS 已分化为四类——传统工程型(MeloTTS)、零样本克隆型(F5-TTS / Spark-TTS / CosyVoice / IndexTTS2)、LLM 化生成型(VoxCPM / Qwen3-TTS / Fish Audio S2 / CosyVoice)、生产服务型(Qwen TTS Realtime)。

其「版本矩阵」表(逐字,截取列)中的首包延迟与许可证:

模型最新版本发布时间许可证商业可用首包延迟
VoxCPM2VoxCPM22026-04Apache-2.0✅RTF ≈ 0.13
Qwen3-TTS1.7B / 0.6B2026-01-22Apache-2.0✅端到端 97ms(首字符即出)
Qwen TTS Realtime云服务持续迭代阿里云商业✅(按量计费)< 200ms(云端 SLA)
CosyVoice 2CosyVoice2-0.5B2024-12Apache-2.0✅首包 150ms
F5-TTS1.x2024MIT✅取决于推理后端
IndexTTS22.02025-09Apache-2.0✅取决于部署
Fish Audio S2 / S2-ProS2 / S2-Pro2026-03-11S2 相对开放 / S2-Pro Research License⚠️ S2-Pro 商用需单独授权< 150ms(官方指标)
Spark-TTS0.5B2025Apache-2.0✅—
MeloTTS持续维护2024-12 末次主提交MIT✅CPU 实时

该文对数字人场景的直接结论(逐字):

第二类是零样本声音克隆型 TTS。代表包括 F5-TTS、Spark-TTS、CosyVoice、IndexTTS2 等。… 这类模型很适合短视频配音、有声书、个人声音助手、数字人和角色语音。

IndexTTS2 … 它的另一个关键能力是时长控制。视频配音和口型同步经常要求一句话必须在固定时间内说完。 … IndexTTS2 的 duration control 对影视、动画、虚拟人、字幕对齐非常有意义。

总体判断:IndexTTS2 适合做情绪配音、角色语音、影视/视频配音、有声书和数字人内容生成。它的上限很高,但商业集成要谨慎。

CosyVoice … 总体判断:CosyVoice 是最适合作为"本地中文 TTS 工程底座"的模型之一。 它未必在每个单项上都第一,但综合能力强,适合认真投入。

总体判断:F5-TTS 适合做轻量自部署、声音克隆、内容生产和实验平台。它是"好用的生成工具",但未必是"实时语音基础设施"的最优解。

Fish Audio S2 Pro 模型卡显示 Research License,非商业免费,商业使用需要单独授权。

补充说明(诚实标注):该文的许可证/延迟数据来自第三方博客转述,本次调研未逐一回溯其原始模型卡;引用时请以该博文为来源。

11.3 与框架原生列表的交叉验证(本次调研自洽性检查)

社区 2026 推荐LiveTalkingOpenAvatarChataigcpanel备注
IndexTTS2✅ indextts2❌✅ IndexTTSLiveTalking 文档给"延时4s / 17G"(OmniTTS 路径),CSDN 实测"首包约1.2s / 8G"
FishSpeech✅ fishtts❌✅ FishSpeech三处一致
CosyVoice✅ cosyvoice✅ 本地 + 百炼✅ 3 个版本覆盖最广
GPT-SoVITS✅ gpt-sovits❌✅—
Edge TTS✅ edgetts(默认)✅❌(不在声音合成表)LiveTalking / OAC 的零成本起步项
Qwen3-TTS / VoxCPM2✅(经 OmniTTS/vLLM-Omni)⚠️ 仅 Qwen-Omni 多模态路径❌LiveTalking 通过统一 OmniTTS 网关接新模型
SparkTTS❌❌✅仅 aigcpanel
F5-TTS / MeloTTS / Bark / XTTSXTTS ✅(xtts,未列入 config 注释)❌❌Open-LLM-VTuber 覆盖 F5?/MeloTTS/Bark

共性结论(2026-09):

  1. CosyVoice + Edge TTS 几乎是所有框架的最小公倍数(LiveTalking / OpenAvatarChat / Linly-Talker / aigcpanel / Open-LLM-VTuber 全部支持)。
  2. GPT-SoVITS / FishSpeech / IndexTTS2 是第二梯队,在 LiveTalking 与 aigcpanel 里都有原生位。
  3. 纯"音频驱动"的说话头项目(SoulX-FlashTalk、Ultralight-Digital-Human、Duix.Avatar 的合成段)自身不带 TTS,选型自由度最高但需自建 TTS 服务;统一落到 16 kHz 单声道(SoulX 明确 sample_rate: 16000;Ultralight 明确"音频采样率需要是16000")。
  4. 实时链路对 TTS 的硬约束是"流式 + 20ms 级 chunk",LiveTalking 是最明确的一例(sample_rate = 16000、chunk = 16000//(fps*2) = 320 samples = 20ms)。

12. 来源清单(本次实际抓取的全部 URL)

12.1 shields.io 元数据

  1. https://img.shields.io/github/stars/lipku/LiveTalking.json
  2. https://img.shields.io/github/license/lipku/LiveTalking.json
  3. https://img.shields.io/github/last-commit/lipku/LiveTalking.json
  4. https://img.shields.io/github/stars/HumanAIGC/OpenAvatarChat.json(返回 repo not found → 用于确认 owner 变更)
  5. https://img.shields.io/github/stars/HumanAIGC-Engineering/OpenAvatarChat.json
  6. https://img.shields.io/github/license/HumanAIGC-Engineering/OpenAvatarChat.json
  7. https://img.shields.io/github/last-commit/HumanAIGC-Engineering/OpenAvatarChat.json
  8. https://img.shields.io/github/stars/xszyou/Fay.json
  9. https://img.shields.io/github/license/xszyou/Fay.json
  10. https://img.shields.io/github/last-commit/xszyou/Fay.json
  11. https://img.shields.io/github/stars/Korvo-AI/Linly-Talker.json(repo not found → owner 变更)
  12. https://img.shields.io/github/stars/Kedreamix/Linly-Talker.json
  13. https://img.shields.io/github/license/Kedreamix/Linly-Talker.json
  14. https://img.shields.io/github/last-commit/Kedreamix/Linly-Talker.json
  15. https://img.shields.io/github/stars/duixcom/Duix.Heygem.json
  16. https://img.shields.io/github/license/duixcom/Duix.Heygem.json
  17. https://img.shields.io/github/last-commit/duixcom/Duix.Heygem.json
  18. https://img.shields.io/github/stars/duixcom/Duix.Avatar.json
  19. https://img.shields.io/github/license/duixcom/Duix.Avatar.json
  20. https://img.shields.io/github/last-commit/duixcom/Duix.Avatar.json
  21. https://img.shields.io/github/stars/modstart-lib/aigcpanel.json
  22. https://img.shields.io/github/license/modstart-lib/aigcpanel.json
  23. https://img.shields.io/github/last-commit/modstart-lib/aigcpanel.json
  24. https://img.shields.io/github/stars/Soul-AILab/SoulX-FlashTalk.json
  25. https://img.shields.io/github/license/Soul-AILab/SoulX-FlashTalk.json
  26. https://img.shields.io/github/last-commit/Soul-AILab/SoulX-FlashTalk.json
  27. https://img.shields.io/github/stars/anliyuan/Ultralight-Digital-Human.json
  28. https://img.shields.io/github/license/anliyuan/Ultralight-Digital-Human.json
  29. https://img.shields.io/github/last-commit/anliyuan/Ultralight-Digital-Human.json
  30. https://img.shields.io/github/stars/Open-LLM-VTuber/Open-LLM-VTuber.json
  31. https://img.shields.io/github/license/Open-LLM-VTuber/Open-LLM-VTuber.json
  32. https://img.shields.io/github/last-commit/Open-LLM-VTuber/Open-LLM-VTuber.json

12.2 README / 源码 / 配置(GitHub raw)

  1. https://raw.githubusercontent.com/lipku/LiveTalking/main/README.md
  2. https://raw.githubusercontent.com/lipku/LiveTalking/main/README-EN.md
  3. https://raw.githubusercontent.com/HumanAIGC-Engineering/OpenAvatarChat/main/README.md
  4. https://raw.githubusercontent.com/HumanAIGC-Engineering/OpenAvatarChat/main/readme_en.md
  5. https://raw.githubusercontent.com/HumanAIGC/OpenAvatarChat/main/README.md(失败:仓库不存在)
  6. https://raw.githubusercontent.com/xszyou/Fay/master/README.md
  7. https://raw.githubusercontent.com/xszyou/Fay/master/system.conf.bak
  8. https://raw.githubusercontent.com/xszyou/Fay/master/tts/ali_tss.py
  9. https://raw.githubusercontent.com/xszyou/Fay/master/tts/ms_tts_sdk.py
  10. https://raw.githubusercontent.com/xszyou/Fay/master/tts/tts_voice.py
  11. https://raw.githubusercontent.com/Kedreamix/Linly-Talker/main/README.md
  12. https://raw.githubusercontent.com/Korvo-AI/Linly-Talker/main/README.md(失败:仓库不存在)
  13. https://raw.githubusercontent.com/duixcom/Duix.Heygem/main/README.md
  14. https://raw.githubusercontent.com/modstart-lib/aigcpanel/main/README.md
  15. https://raw.githubusercontent.com/Soul-AILab/SoulX-FlashTalk/main/README.md
  16. https://raw.githubusercontent.com/Open-LLM-VTuber/Open-LLM-VTuber/main/README.md
  17. https://raw.githubusercontent.com/anliyuan/Ultralight-Digital-Human/main/README.md(404 → 默认分支为 master)

12.3 jsDelivr(CDN 与文件树 API)

  1. https://data.jsdelivr.com/v1/packages/gh/lipku/LiveTalking@main
  2. https://data.jsdelivr.com/v1/packages/gh/HumanAIGC-Engineering/OpenAvatarChat@main
  3. https://data.jsdelivr.com/v1/packages/gh/Soul-AILab/SoulX-FlashTalk@main
  4. https://data.jsdelivr.com/v1/packages/gh/modstart-lib/aigcpanel@master
  5. https://data.jsdelivr.com/v1/packages/gh/Kedreamix/Linly-Talker@main(403)
  6. https://data.jsdelivr.com/v1/packages/gh/xszyou/Fay@master(403)
  7. https://data.jsdelivr.com/v1/packages/gh/duixcom/Duix.Heygem@main(403)
  8. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/config.yaml
  9. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/config.py
  10. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/registry.py
  11. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/requirements.txt
  12. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/__init__.py
  13. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/base_tts.py
  14. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/edge.py
  15. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/azure.py
  16. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/doubao.py
  17. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/omnitts.py
  18. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/qwentts.py
  19. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/sovits.py
  20. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/tencent.py
  21. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/xtts.py
  22. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/cosyvoice.py(404)
  23. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/fish.py(404)
  24. https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/indextts2.py(404)
  25. https://cdn.jsdelivr.net/gh/HumanAIGC-Engineering/OpenAvatarChat@main/src/handlers/tts/cosyvoice/tts_handler_cosyvoice.py
  26. https://cdn.jsdelivr.net/gh/HumanAIGC-Engineering/OpenAvatarChat@main/src/handlers/tts/edgetts/tts_handler_edgetts.py
  27. https://cdn.jsdelivr.net/gh/HumanAIGC-Engineering/OpenAvatarChat@main/config/chat_with_openai_compatible_edge_tts.yaml
  28. https://cdn.jsdelivr.net/gh/xszyou/Fay@master/tts/ali_tss.py
  29. https://cdn.jsdelivr.net/gh/xszyou/Fay@master/tts/ms_tts_sdk.py
  30. https://cdn.jsdelivr.net/gh/xszyou/Fay@master/tts/tts_voice.py
  31. https://cdn.jsdelivr.net/gh/Kedreamix/Linly-Talker@main/TTS/README.md
  32. https://cdn.jsdelivr.net/gh/anliyuan/Ultralight-Digital-Human@master/README.md
  33. https://cdn.jsdelivr.net/gh/Soul-AILab/SoulX-FlashTalk@main/flash_talk/configs/infer_params.yaml
  34. https://cdn.jsdelivr.net/gh/Soul-AILab/SoulX-FlashTalk@main/generate_video.py
  35. https://cdn.jsdelivr.net/gh/weihaox/awesome-digital-human@main/README.md
  36. https://cdn.jsdelivr.net/gh/icemaple77/digital-human@main/README.md

12.4 官方文档站

  1. https://doc.livetalking.ai/docs/tts/
  2. https://doc.livetalking.ai/docs/tts/gptsovits/
  3. https://doc.livetalking.ai/docs/tts/fishspeech/
  4. https://doc.livetalking.ai/docs/tts/cosyvoice/
  5. https://doc.livetalking.ai/docs/tts/omnitts/
  6. https://doc.livetalking.ai/docs/faq/
  7. https://humanaigc-engineering.github.io/OpenAvatarChat/reference/preset-modes
  8. https://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/bailian-cosyvoice.html
  9. https://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/cosyvoice-local.html
  10. https://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/edge-tts.html

12.5 DeepWiki(第三方代码索引,索引时间 2026-03-29)

  1. https://deepwiki.com/xszyou/Fay/4.2-text-to-speech-(tts)
  2. https://deepwiki.com/xszyou/Fay/10.1-local-setup-and-configuration
  3. https://deepwiki.com/xszyou/Fay
  4. https://deepwiki.com/xszyou/Fay/4-inputoutput-systems

12.6 2026 社区 / 博客

  1. https://juejin.cn/post/7649764223339495475 —《开源 TTS 模型横向对比:从"能发声"到"可部署的语音智能基础设施"(2026 版)》,2026-06-12
  2. https://blog.csdn.net/jacke121/article/details/157323875 —《LiveTalking 部署笔记》,2026-01-24 首发 / 2026-03-30 更新

12.7 未查到项汇总(诚实标注)

框架未查到内容
LiveTalkingTTS 级 TTFB 硬阈值(FAQ 未涉及 TTS);cosyvoice.py / fish.py / indextts2.py 的注册名原文(CDN 404)
OpenAvatarChatTTS 级延迟数字(只有系统级"平均响应时间 2.2 秒");TTS chunk 大小约束
FayTTS 级 TTFB / ms 指标;火山引擎(volcano)输出的采样率(README 与源码未在本轮取到该实现细节)
Linly-TalkerTTS 输出采样率 / chunk 约束(README 与 TTS/README.md 均未声明)
Duix.Avatar采样率要求;TTFB(离线任务式);TTS 替换机制文档
aigcpanel任何延迟 / FPS / 流式指标;采样率 / chunk 约束;自定义 TTS 插件文档
SoulX-FlashTalkTTS 相关项不适用(无 TTS)
Ultralight-Digital-HumanTTS 相关项不适用(无 TTS);每帧音频 chunk 大小
Open-LLM-VTuberTTS 级延迟数字;采样率约束
awesome 清单weihaox/awesome-digital-human 不列任何 TTS 后端

报告结束。所有引用均来自上列 URL 的实际抓取内容;凡 jsDelivr 与 raw 存在分歧处(Linly-Talker / Ultralight / OpenAvatarChat 的文件树 vs CDN),已在对应章节显式标注。

下载此文件