调研日期:2026-09-22 调研方法:全部结论来自实际 curl 抓取的原始文件(GitHub raw / jsDelivr / 官方文档站 / 官方配置与源码),逐条引用原文。星标 / 许可证 / 最后提交通过 shields.io JSON 获取(GitHub API 限流,未使用)。 重要说明:
- 网络过程中
raw.githubusercontent.com与github.com多次超时中断,因此对长文件同时使用了cdn.jsdelivr.net与data.jsdelivr.com(jsDelivr 文件树 API)。凡使用 jsDelivr 的位置均已标注。- 凡无法确证的项一律写 未查到,不做推测。
- 本文件所有 URL 均为本次实际成功抓取的地址。
| 框架 | 仓库 | 星标 | 许可证 | 最后提交 | 原生 TTS 引擎 |
|---|---|---|---|---|---|
| LiveTalking | lipku/LiveTalking | 9.6k | Apache-2.0 | september | edgetts / gpt-sovits / cosyvoice / fishtts / tencent / doubao / indextts2 / azuretts / qwentts(config.yaml 原注释);代码另有 xtts、omnitts |
| OpenAvatarChat | HumanAIGC-Engineering/OpenAvatarChat | 3.8k | Apache-2.0 | july | EdgeTTS / CosyVoice(本地)/ 百炼 CosyVoice(API);Qwen-Omni 模式下 TTS 由 Qwen-Omni 本体承担 |
| Fay | xszyou/Fay | 14k | GPL-3.0 | yesterday | azure / ali / gptsovits / volcano / gptsovits_v3 |
| Linly-Talker | Kedreamix/Linly-Talker | 3.5k | MIT | february | Edge TTS / PaddleTTS / GPT-SoVITS(推荐)/ XTTS / CosyVoice |
| Duix.Heygem(现 Duix.Avatar) | duixcom/Duix.Heygem → duixcom/Duix.Avatar | 16k | not identifiable by github | april | fish-speech-ziming(guiji2025/fish-speech-ziming Docker 镜像),固定内置、无切换项 |
| aigcpanel | modstart-lib/aigcpanel | 5.6k | Apache-2.0 | last friday | CosyVoice-300M / CosyVoice-300M-Instruct / CosyVoice2-0.5b / FishSpeech / IndexTTS / SparkTTS / GPT-SoVITS |
| SoulX-FlashTalk | Soul-AILab/SoulX-FlashTalk | 1.5k | Apache-2.0 | july | 无(纯音频驱动,不含任何 TTS) |
| Ultralight-Digital-Human | anliyuan/Ultralight-Digital-Human | 2.6k | not specified | july | 无(纯音频特征驱动,不含任何 TTS) |
| Open-LLM-VTuber(参考) | Open-LLM-VTuber/Open-LLM-VTuber | 14k | not identifiable by github | may | sherpa-onnx / pyttsx3 / MeloTTS / Coqui-TTS / GPTSoVITS / Bark / CosyVoice / Edge TTS / Fish Audio / Azure TTS 等 |
| Awesome 类清单 | weihaox/awesome-digital-human | — | — | — | 未查到(该清单聚焦 3D 数字人/avatar 生成,正文未列 TTS 后端) |
| Awesome 类清单 | icemaple77/digital-human | — | — | — | edge(免费)/ GPT-SoVITS(自训音色) |
| 项 | 值 |
|---|---|
| 仓库 | https://github.com/lipku/LiveTalking |
| 星标 | {"label":"stars","message":"9.6k",...} |
| 许可证 | {"label":"license","message":"Apache-2.0",...} |
| 最后提交 | {"label":"last commit","message":"september",...} |
抓取地址:
https://img.shields.io/github/stars/lipku/LiveTalking.jsonhttps://img.shields.io/github/license/lipku/LiveTalking.jsonhttps://img.shields.io/github/last-commit/lipku/LiveTalking.json来源:https://raw.githubusercontent.com/lipku/LiveTalking/main/README.md(另抓 README-EN.md,同一提交)
"3. 系统架构 / 逻辑层":
- TTS 引擎: 模块化设计,支持 EdgeTTS、GPT-SoVITS、CosyVoice、腾讯云等多种方案
英文版同源文件 https://raw.githubusercontent.com/lipku/LiveTalking/main/README-EN.md 第 142 行逐字对应(英文句末还有 "and more"):
- TTS Engine: Modular design supporting EdgeTTS, GPT-SoVITS, CosyVoice, Tencent Cloud, and more
"插件系统":
- 基于 registry.py 的去中心化注册机制,开发者可自行扩展 TTS、Avatar、Output 模块
Features:
- 支持多种数字人模型: ernerf、musetalk、wav2lip、Ultralight-Digital-Human
- 支持声音克隆
- 支持数字人说话被打断
核心流程:
核心流程:用户输入文字/音频 → LLM 生成回复(可选)→ TTS 合成语音 → 数字人实时口型同步 → 音视频推流输出
config.yaml(源码级,最可靠)来源:https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/config.yaml (等价 raw:https://raw.githubusercontent.com/lipku/LiveTalking/main/config.yaml)
# -------- TTS ----------------------------------------------------------
tts: edgetts # edgetts / gpt-sovits / cosyvoice / fishtts / tencent / doubao / indextts2 / azuretts / qwentts
REF_FILE: '' # voice name or reference audio file path (wav, 16kHz, mono)
REF_TEXT: '' # reference text for voice cloning (English or Chinese)
TTS_SERVER: '' # TTS server URL, no trailing slash
→ 文档化的官方 TTS 枚举(9 个):edgetts、gpt-sovits、cosyvoice、fishtts、tencent、doubao、indextts2、azuretts、qwentts。
文件树来源:https://data.jsdelivr.com/v1/packages/gh/lipku/LiveTalking@main
/tts/__init__.py
/tts/azure.py
/tts/base_tts.py
/tts/cosyvoice.py
/tts/doubao.py
/tts/edge.py
/tts/fish.py
/tts/indextts2.py
/tts/omnitts.py
/tts/qwentts.py
/tts/sovits.py
/tts/tencent.py
/tts/xtts.py
已从源码中逐一核实的注册名(@register("tts", "<name>") 装饰器):
| 文件 | 注册名 | 抓取地址(jsDelivr CDN) |
|---|---|---|
tts/edge.py | edgetts | https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/edge.py |
tts/azure.py | azuretts | .../tts/azure.py |
tts/doubao.py | doubao | .../tts/doubao.py |
tts/omnitts.py | omnitts | .../tts/omnitts.py |
tts/qwentts.py | qwentts | .../tts/qwentts.py |
tts/sovits.py | gpt-sovits | .../tts/sovits.py |
tts/tencent.py | tencent | .../tts/tencent.py |
tts/xtts.py | xtts | .../tts/xtts.py |
tts/cosyvoice.py | 未取到(CDN 404,但文件树中存在) | .../tts/cosyvoice.py |
tts/fish.py | 未取到(同上) | .../tts/fish.py |
tts/indextts2.py | 未取到(同上) | .../tts/indextts2.py |
说明:
cosyvoice.py/fish.py/indextts2.py三个文件在 jsDelivr 文件树 API 中存在,但 CDN 直取返回 404(jsDelivr 的 tree 与 CDN 可能解析到不同提交),因此其注册名未能直接引用原文;对应的 config 名(cosyvoice/fishtts/indextts2)来自上面config.yaml的原文注释。 另注:xtts与omnitts存在于代码,但未出现在config.yaml的注释枚举里。
来源:https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/registry.py
_REGISTRY: Dict[str, Dict[str, Type]] = {
"stt": {}, "llm": {}, "tts": {}, "avatar": {}, "output": {},
}
def register(category: str, name: str):
"""
装饰器:注册插件类到全局注册表。
用法::
@register("tts", "edgetts")
class EdgeTTS(BaseTTS): ...
"""
以及 create(category, name, **kwargs) 的报错文案原文:
f"Plugin '{name}' not found in category '{category}'. Available: {available}"
→ 结论:TTS 层完全可插拔。任意自定义 TTS 只需继承 BaseTTS 并加 @register("tts", "<name>"),再用 tts: <name> 启用。
stream=True 并逐 chunk 推送,且显式打印首包延迟。逐字引用(各文件):
tts/sovits.py:logger.info(f"gpt_sovits Time to first chunk: {end-start}s"),请求体 'streaming_mode':True;注释原文 # #req["stream_chunk_size"] = stream_chunk_size # you can reduce it to get faster response, but degrade qualitytts/xtts.py:"20" #args.stream_chunk_size,注释 # you can reduce it to get faster response, but degrade qualitytts/tencent.py:for chunk in res.iter_content(chunk_size=6400): # 640 16K*20ms*2tts/sovits.py:for chunk in res.iter_content(chunk_size=None): #12800 1280 32K*20ms*2tts/xtts.py:for chunk in res.iter_content(chunk_size=None): #24K*20ms*2
- 后端日志
inferfps= GPU 推理帧率,finalfps= 最终推流帧率,两者均需 >=25 才算实时
来源:https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/base_tts.py
class BaseTTS:
def __init__(self, opt, parent: "BaseAvatar"):
self.opt = opt
self.parent = parent
#self.fps = opt.fps # 20 ms per frame
self.sample_rate = 16000
self.chunk = self.sample_rate // (opt.fps*2) # 320 samples per chunk (20ms * 16000 / 1000)
16000 // (fps*2) 个采样点 = 默认 fps=25 时 320 samples = 20 msconfig.yaml 原文:REF_FILE: '' # voice name or reference audio file path (wav, 16kHz, mono)tts/qwentts.py 亦显式以 16k 对齐:sample_rate=16000、注释 # 按 self.chunk (320 samples = 20ms @16kHz) 分块推送tts/azure.py 以字节算 chunk:chunk_bytes = self.chunk * 2 # 320 samples * 2 bytes (int16)来源:https://doc.livetalking.ai/docs/tts/
页面为 TTS 索引,正文逐字:
TTS 相关文档索引。 GPT-SoVITS Fish Speech CosyVoice 旧版说明
已抓取的子页:
| 子页 | URL | 要点(逐字) |
|---|---|---|
| GPT-SoVITS | https://doc.livetalking.ai/docs/tts/gptsovits/ | "采用gpt-sovits方案,bert-sovits适合长音频训练,gpt-sovits运行短音频快速推理";endpoint: /tts;示例含 streaming_mode=true;"media_type" : "wav" # support "wav" , "raw" , "ogg" , "aac" |
| Fish Speech | https://doc.livetalking.ai/docs/tts/fishspeech/ | endpoint: /v1/tts;"chunk_length" : 200、"format" : "wav"、"streaming" : false(默认)、fishaudio/fish-speech-1.5 |
| CosyVoice | https://doc.livetalking.ai/docs/tts/cosyvoice/ | 用 iic/CosyVoice2-0.5B + iic/CosyVoice-ttsfrd;启动 CosyVoice/runtime/python/fastapi/server.py --model_dir .../CosyVoice2-0.5B |
| OmniTTS | https://doc.livetalking.ai/docs/tts/omnitts/ | 一个 OpenAI 兼容的统一 TTS 服务(vLLM-Omni),覆盖多个模型——见下 |
OmniTTS 页逐字要点(这是 LiveTalking 官方"一站式接多 TTS"的答案):
vLLM-Omni是一个推理服务系统,已经支持了多种tts模型,包括qwen3-tts、cosyvoice、VoxCpm2。更多模型 https://docs.vllm.ai/projects/vllm-omni/en/latest/serving/speech_api/ 。对外提供统一的tts API接口服务,兼容openai规范,支持克隆声音的上传、删除、查看。
其 "3. 启动api服务" 小节给出的模型与显存/延迟(逐字摘录):
| 小节 | 命令 | 原文备注 |
|---|---|---|
| 3.1 Qwen3 | vllm serve Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice --omni --trust-remote-code --port 8091 | "需要显存12g" |
| 3.2 Voxcpm | vllm serve openbmb/VoxCPM2 --omni --trust-remote-code --port 8091 | "需要显存13G" |
| 3.3 Fish Speech | vllm serve fishaudio/s2-pro --omni --trust-remote-code --port 8091 | "需要显存18G" |
| 3.4 IndexTTS | vllm serve IndexTeam/IndexTTS-2 --omni --trust-remote-code --port 8091 | "需要显存17G,延时4s" |
| 3.5 CosyVoice | (uv pip install s3tokenizer …) | "需要显存14G" |
https://doc.livetalking.ai/docs/faq/ 抓取成功,但正文仅见 pytorch3d 安装问题,未涉及 TTS 延迟或格式 → TTS 相关 Q&A 未查到。
⚠️ 重要更正:任务给出的
HumanAIGC/OpenAvatarChat不存在(shields.io 返回"message":"repo not found")。实际仓库所有者为HumanAIGC-Engineering:https://github.com/HumanAIGC-Engineering/OpenAvatarChat。 抓取https://raw.githubusercontent.com/HumanAIGC/OpenAvatarChat/main/README.md亦失败(非 200)。
| 项 | 值 |
|---|---|
| 仓库 | https://github.com/HumanAIGC-Engineering/OpenAvatarChat |
| 星标 | {"label":"stars","message":"3.8k",...} |
| 许可证 | {"label":"license","message":"Apache-2.0",...} |
| 最后提交 | {"label":"last commit","message":"july",...} |
抓取地址:https://img.shields.io/github/{stars,license,last-commit}/HumanAIGC-Engineering/OpenAvatarChat.json
来源:https://raw.githubusercontent.com/HumanAIGC-Engineering/OpenAvatarChat/main/README.md
核心亮点:
- 模块化架构设计:采用高度模块化设计,可灵活替换 ASR、LLM、TTS、Avatar 等核心组件
- 低延迟优化:通过 VAD 检测、语音缓冲、帧率控制等机制优化,平均响应时间仅 2.2 秒
组件依赖表(TTS 行):
| TTS | FunAudioLLM/CosyVoice |[GitHub]](https://github.com/FunAudioLLM/CosyVoice)||
预置模式表(TTS 列):
| CONFIG名称 | ASR | LLM | TTS | AVATAR | | chat_with_lam.yaml | SenseVoice | API | API | LAM | | chat_with_qwen_omni.yaml | Qwen-Omni | Qwen-Omni | Qwen-Omni | lite-avatar | | chat_with_openai_compatible_bailian_cosyvoice.yaml | SenseVoice | API | API | lite-avatar | | chat_with_openai_compatible_bailian_cosyvoice_flashhead.yaml | SenseVoice | API | API | FlashHead |
最新动态(2026.04,v0.6.0):
- [2026.04] ⭐️⭐️⭐️ 版本 0.6.0发布: … 接入 SoulX-FlashHead 数字人,基于扩散模型的实时流式说话头生成
来源:https://humanaigc-engineering.github.io/OpenAvatarChat/reference/preset-modes
侧边栏 "TTS(语音合成)" 下只有三项,逐字:
TTS(语音合成) 百炼 CosyVoice CosyVoice 本地 Edge TTS
对应文档页(均抓取成功):
| Handler | 文档 URL | 关键点(逐字) |
|---|---|---|
| 百炼 CosyVoice | https://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/bailian-cosyvoice.html | model_name : "cosyvoice-v1";TTS_CosyVoice.sample_rate | 24000 | 输出音频采样率 |
| CosyVoice 本地 | https://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/cosyvoice-local.html | "TTS 默认为 CosyVoice 的 iic/CosyVoice-300M-SFT + 中文女,可以通过修改为其他模型配合 ref_audio_path 和 ref_audio_text 进行音色复刻。" |
| Edge TTS | https://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/edge-tts.html | "集成微软的 Edge TTS,使用云端推理,无需申请 API Key。" 配置:Edge_TTS: module: tts/edgetts/tts_handler_edgetts / voice: "zh-CN-XiaoxiaoNeural" |
文件树来源:https://data.jsdelivr.com/v1/packages/gh/HumanAIGC-Engineering/OpenAvatarChat@main
/src/handlers/tts/bailian_tts/tts_handler_cosyvoice_bailian.py
/src/handlers/tts/cosyvoice/tts_handler_cosyvoice.py
/src/handlers/tts/cosyvoice/cosyvoice_processor.py
/src/handlers/tts/edgetts/tts_handler_edgetts.py
/config/chat_with_openai_compatible_edge_tts.yaml
/config/chat_with_openai_compatible_bailian_cosyvoice.yaml
/config/chat_with_gs.yaml
→ 确认 OpenAvatarChat 原生 TTS 只有 3 条路径:本地 CosyVoice、百炼 CosyVoice API、Edge TTS。(chat_with_qwen_omni.yaml 模式下 TTS 由 Qwen-Omni 多模态模型本体输出,不算独立 TTS 引擎。)
来源:https://cdn.jsdelivr.net/gh/HumanAIGC-Engineering/OpenAvatarChat@main/config/chat_with_openai_compatible_edge_tts.yaml
chat_engine:
model_root: "models"
concurrent_limit: 1
handler_search_path:
- "src/handlers"
handler_configs:
Edge_TTS:
enabled: True
module: tts/edgetts/tts_handler_edgetts
voice: "zh-CN-XiaoxiaoNeural"
→ 通过 handler_search_path + module: 声明式装配,新增 TTS 只要按 HandlerBase 写一个模块即可。(框架无"任意 TTS 免代码接入"的网关,需要写 handler。)
平均响应时间仅 2.2 秒
tts_handler_cosyvoice.py):inputs.is_last_data 为假时按标点切句并逐句入队——
sentences = re.split(r'(?<=[,.~!?,。!?])', context.input_text)
if len(sentences) > 1: # 至少有一个完整句子
句尾会插入一段静音结束标记:end_task.result_queue.put(np.zeros(shape=(1, 240), dtype=np.float32))
来源:https://cdn.jsdelivr.net/gh/HumanAIGC-Engineering/OpenAvatarChat@main/src/handlers/tts/cosyvoice/tts_handler_cosyvoice.py 与 .../edgetts/tts_handler_edgetts.py
两个 Handler 的 TTSConfig 都声明:
sample_rate: int = Field(default=24000)
sample_rate: int = Field(default=24000)
输出为单声道(DataBundleEntry.create_audio_entry("avatar_audio", 1, self.sample_rate) 的第 2 个参数 1 = 通道数):
definition.add_entry(DataBundleEntry.create_audio_entry("avatar_audio", 1, self.sample_rate))
sample_rate 可在配置中改(百炼文档亦写 TTS_CosyVoice.sample_rate | 24000 | 输出音频采样率)。fps: 25)。| 项 | 值 |
|---|---|
| 仓库 | https://github.com/xszyou/Fay |
| 星标 | {"label":"stars","message":"14k",...} |
| 许可证 | {"label":"license","message":"GPL-3.0",...} |
| 最后提交 | {"label":"last commit","message":"yesterday",...} |
抓取地址:https://img.shields.io/github/{stars,license,last-commit}/xszyou/Fay.json
来源:https://raw.githubusercontent.com/xszyou/Fay/master/README.md(注意默认分支为 master)
Fay数字人框架,向上适配各种数字人模型技术,向下接入各式大语言模型,并且便于更换诸如TTS、ASR等模型,为单片机、app、网站提供全面的数字人应用接口。
功能特点:
- 自由匹配数字人模型、大语言模型(openai 兼容接口)、ASR、TTS模型
- 全时流式的支持
- 支持唤醒及打断对话
⚠️ README 正文没有列出具体 TTS 引擎名称,只给飞书文档链接。因此下面的引擎清单来自源码配置(更权威)。
system.conf.bak(逐字引用)来源:https://raw.githubusercontent.com/xszyou/Fay/master/system.conf.bak
#tts类型(切换请重新选择所需要的声音)azure、ali、gptsovits、volcano、gptsovits_v3
tts_module=ali
# 微软 文字转语音 服务密钥(非必须,使用可产生不同情绪的音频)https://azure.microsoft.com/zh-cn/services/cognitive-services/text-to-speech/
ms_tts_key=
ms_tts_region=
# 阿里云 文字转语音 服务密钥 https://ai.aliyun.com/nls/trans
ali_tss_key_id=
ali_tss_key_secret=
ali_tss_app_key=
# Doubao-语音合成 服务密钥 https://www.volcengine.com/product/voice-tech
volcano_tts_appid=
volcano_tts_access_token=
volcano_tts_cluster=volcano_tts
#可为空,为空时读取选择的音色
volcano_tts_voice_type=
→ Fay 官方 TTS 引擎(5 个):azure(微软)、ali(阿里云 NLS)、gptsovits、volcano(火山引擎/豆包语音)、gptsovits_v3。
旁证(DeepWiki,索引时间 "Last indexed: 29 March 2026"):https://deepwiki.com/xszyou/Fay/10.1-local-setup-and-configuration
| tts_module | Text-to-speech engine | ali , azure , volcano , gptsovits |
(DeepWiki 该表未列 gptsovits_v3;system.conf.bak 为更权威来源。)
来源:https://raw.githubusercontent.com/xszyou/Fay/master/tts/ali_tss.py、.../tts/ms_tts_sdk.py、.../tts/tts_voice.py (三者与 jsDelivr 版本 https://cdn.jsdelivr.net/gh/xszyou/Fay@master/tts/* 内容一致,已逐行 diff 一致)
tts/ali_tss.py — 阿里云 NLS class Speech,含 token 管理与历史缓存tts/ms_tts_sdk.py — 微软 Azure + Edge 双路径:import azure.cognitiveservices.speech as speechsdk
import edge_tts
...
if config_util.key_ms_tts_key and config_util.key_ms_tts_key is not None and config_util.key_ms_tts_key.strip() != "":
self.__speech_config = speechsdk.SpeechConfig(subscription=cfg.key_ms_tts_key, region=cfg.key_ms_tts_region)
...
self.__speech_config.set_speech_synthesis_output_format(speechsdk.SpeechSynthesisOutputFormat.Riff16Khz16BitMonoPcm)
self.__synthesizer = speechsdk.SpeechSynthesizer(speech_config=self.__speech_config, audio_config=None)
self.ms_tts = True
tts/tts_voice.py — 音色/情绪注册表 EnumVoice,例如:class EnumVoice(Enum):
XIAO_XIAO_NEW = {
"name": "晓晓(azure)",
"voiceName": "zh-CN-XiaoxiaoMultilingualNeural",
"styleList": {"angry": "angry", "lyrical": "lyrical", "calm": "gentle",
"assistant": "affectionate", "cheerful": "cheerful"}
}
XIAO_XIAO = { "name": "晓晓(edge)", "voiceName": "zh-CN-XiaoxiaoNeural", ... }
DeepWiki TTS 页(https://deepwiki.com/xszyou/Fay/4.2-text-to-speech-(tts))补充(逐字):
Azure SDK : If key_ms_tts_key is provided, it uses the official speechsdk for high-fidelity synthesis with SSML support for fine-grained style control Edge TTS (Fallback) : If no API key is found, it falls back to the edge_tts library, which provides free access to Microsoft's neural voices via the Edge browser's API Format Conversion : Edge TTS outputs .mp3 , which is automatically converted to .wav using pydub to ensure compatibility with the Fay audio player Audio Processing : Converts the returned stream into a WAV file with 16kHz sampling rate and mono channel
system.conf 的 tts_module(原文注释:#tts类型(切换请重新选择所需要的声音)…)。只能选上述 5 个内置实现;要接别的 TTS 需自行改 tts/ 下的代码。tts/tts_voice.py 的中央注册表 + get_voice_of(name) 管理(原文:If a configured voice is not found, the system defaults to "晓晓(edge)")。
- 全时流式的支持
- 支持唤醒及打断对话
In streaming mode (e.g., when using LLMs), the TTS system must coordinate with the UI to handle audio playback timing. The StreamStateManager manages hidden markers appended to the text. The markers
<isfirst>and<isend>are used by the WebSocket server to tell the frontend whether to clear the audio queue or finish the animation
Riff16Khz16BitMonoPcm(源码硬编码,见上).mp3,经 pydub 转为 .wav(DeepWiki 原文)⚠️ 重要更正:任务给出的
Korvo-AI/Linly-Talker不存在(shields.io 返回"message":"repo not found")。实际仓库为Kedreamix/Linly-Talker。
| 项 | 值 |
|---|---|
| 仓库 | https://github.com/Kedreamix/Linly-Talker |
| 星标 | {"label":"stars","message":"3.5k",...} |
| 许可证 | {"label":"license","message":"MIT",...} |
| 最后提交 | {"label":"last commit","message":"february",...} |
抓取地址:https://img.shields.io/github/{stars,license,last-commit}/Kedreamix/Linly-Talker.json
来源:https://raw.githubusercontent.com/Kedreamix/Linly-Talker/main/README.md
README 目录树原文(第 87–99 行):
- [ASR - Speech Recognition](#asr---speech-recognition)
- [Whisper](#whisper)
- [FunASR](#funasr)
- [Coming Soon](#coming-soon)
- [TTS - Text To Speech](#tts---text-to-speech)
- [Edge TTS](#edge-tts)
- [PaddleTTS](#paddletts)
- [Coming Soon](#coming-soon-1)
- [Voice Clone](#voice-clone)
- [GPT-SoVITS(Recommend)](#gpt-sovitsrecommend)
- [XTTS](#xtts)
- [CosyVoice](#cosyvoice)
- [Coming Soon](#coming-soon-2)
→ 文档化的 TTS / 声音克隆后端(5 个):Edge TTS、PaddleTTS、GPT-SoVITS(标注 Recommend)、XTTS、CosyVoice。 (README 把前两者归在 "TTS",后三者归在 "Voice Clone",但四者都是语音合成实现。)
关键原文摘录:
Edge TTS
To use Microsoft Edge's online text-to-speech service from Python without needing Microsoft Edge or Windows or an API key, you can refer to ... https://github.com/rany2/edge-tts
Due to some issues with the Edge TTS repository, it seems that Microsoft has restricted certain IPs. ... I recommend using the CosyVoice method.
PaddleTTS
In practical use, there may be scenarios that require offline operation. Since Edge TTS requires an online environment to generate speech, we have chosen PaddleSpeech, another open-source alternative, for Text-to-Speech (TTS). ... https://github.com/PaddlePaddle/PaddleSpeech
GPT-SoVITS(Recommend)
Thank you for your open source contribution. I have also found the
GPT-SoVITSvoice cloning model to be quite impressive. ... https://github.com/RVC-Boss/GPT-SoVITSXTTS
Coqui XTTS is a leading deep learning toolkit for Text-to-Speech (TTS) tasks, allowing for voice cloning and voice transfer to different languages using a 5-second or longer audio clip.
CosyVoice
CosyVoice is an open-source multilingual speech understanding model developed by Alibaba's Tongyi Lab ... 1. CosyVoice-300M ... 2. CosyVoice-300M-SFT ... 3. CosyVoice-300M-Instruct ...
更新日志中的相关条目:
- **Updated the offline mode for Paddle TTS, excluding Edge TTS.**
- **Implemented a simple fix for the Edge-TTS bug, resolved several issues with MuseTalk, and plan to integrate fishTTS for more stable TTS performance**
("plan to integrate fishTTS" —— 计划中,非已完成。)
架构定位(原文):
- 🧩 Modular Multimodal Pipeline: Reuses existing ASR/LLM/TTS/Avatar capabilities while adopting streaming processing framework for system refactoring.
TTS/README.md来源:https://cdn.jsdelivr.net/gh/Kedreamix/Linly-Talker@main/TTS/README.md(3433 B,同时以 raw 路径核对过)
标题逐字:
TTS 赋予数字人真实的语音交互能力
Edge-TTS
PaddleTTS
正文包含自实现的 EdgeTTS 包装类(SUPPORTED_VOICE、predict(...) 输出 result.wav + result.vtt 字幕)与 PaddleSpeech 的 TTSExecutor 封装(am='fastspeech2'/'tacotron2'、voc='pwgan'/'hifigan'/...、lang='zh'/'en'/'mix'/'canton')。
TTS/README.md 均未声明 TTS 采样率 / chunk 约束。→ 未查到(需从其 SadTalker / Wav2Lip 的音频预处理代码反推,本次未取到对应源码文件)。| 项 | Duix.Heygem | Duix.Avatar |
|---|---|---|
| 星标 | {"label":"stars","message":"16k",...} | {"label":"stars","message":"16k",...} |
| 许可证 | {"label":"license","message":"not identifiable by github",...} | 同左 |
| 最后提交 | {"label":"last commit","message":"april",...} | {"label":"last commit","message":"april",...} |
抓取地址:https://img.shields.io/github/{stars,license,last-commit}/duixcom/Duix.Heygem.json 与 .../duixcom/Duix.Avatar.json
→ 两仓库三项指标完全一致,且 README 内容即 "Duix.Avatar";README 内 LICENSE 链接指向 https://github.com/duixcom/Duix.Avatar/blob/main/LICENSE。结论:Duix.Heygem 已改名为 Duix.Avatar(同一仓库)。
来源:https://raw.githubusercontent.com/duixcom/Duix.Heygem/main/README.md
依赖(Docker 镜像):
- Nodejs 18
- Docker Images
- docker pull guiji2025/fun-asr
- docker pull guiji2025/fish-speech-ziming
致谢章节(最直接的一句):
10. Acknowledgments
- ASR based on fun-asr
- TTS based on fish-speech-ziming
关键路径:
- src/main/service/voice.js
- Separate video into silent video + audio
- Place audio in
D:\duix_avatar_data\voice\datais agreed with theguiji2025/fish-speech-zimingservice, can be modified in docker-compose
→ 原生 TTS 只有 1 个:fish-speech-ziming(GuiJi/硅基智能定制的 Fish-Speech Docker 服务,guiji2025/fish-speech-ziming)。 无其他内置 TTS。
docker-compose 中修改路径,"可修改和扩展代码"是泛泛的本地部署对比表用语,没有 TTS 插件/注册表机制。Duix.Avatar is a fully offline video synthesis tool designed for Windows systems
"streaming": false, // Fixed parameter —— 明确关闭流式(逐字引用):### **Audio Synthesis**
Interface: `http://127.0.0.1:18180/v1/invoke`
{
"speaker": "{uuid}", // A unique UUID
"text": "xxxxxxxxxx", // Text content to synthesize
"format": "wav", // Fixed parameter
"topP": 0.7, // Fixed parameter
"max_new_tokens": 1024, // Fixed parameter
"chunk_length": 100, // Fixed parameter
"repetition_penalty": 1.2, // Fixed parameter
"temperature": 0.7, // Fixed parameter
"need_asr": false, // Fixed parameter
"streaming": false, // Fixed parameter
"is_fixed_seed": 0, // Fixed parameter
"is_norm": 0, // Fixed parameter
"reference_audio": "{voice.asr_format_audio_url}", // Return value from previous "Model Training" step
"reference_text": "{voice.reference_audio_text}" // Return value from previous "Model Training" step
}
"format": "wav" —— README 标注 // Fixed parameter(固定为 wav)。"chunk_length": 100 —— 同样是固定参数(属于 fish-speech 的 chunk_length)。http://127.0.0.1:8383/easy/submit,参数 audio_url / video_url / code / chaofen / watermark_switch / pn。)owner 确认:正确仓库为
modstart-lib/aigcpanel(抓取成功);README 亦引用https://aigcpanel.com与 Gitee 同路径。
| 项 | 值 |
|---|---|
| 仓库 | https://github.com/modstart-lib/aigcpanel |
| 星标 | {"label":"stars","message":"5.6k",...} |
| 许可证 | {"label":"license","message":"Apache-2.0",...} |
| 最后提交 | {"label":"last commit","message":"last friday",...} |
抓取地址:https://img.shields.io/github/{stars,license,last-commit}/modstart-lib/aigcpanel.json
来源:https://raw.githubusercontent.com/modstart-lib/aigcpanel/main/README.md
软件介绍:
AIGCPanel 是一款简单易用的一站式 AI 数字人桌面应用,支持 Windows / macOS / Linux 三平台。… 核心能力涵盖:数字人视频合成(换口型)、语音合成 / 克隆 / 识别、25+ 音视频处理工具、智能直播互动,Pro 版额外提供可视化工作流编排和云端 AI 模型服务。 软件内置模型市场,支持一键下载启动包,开箱即用;同时兼容远程 API 模型,灵活适配各类部署场景。
"支持模型 → 声音合成"表格(逐字):
声音合成
| 模型 | 说明 | | CosyVoice-300M | 阿里通义实验室开源 TTS | | CosyVoice-300M-Instruct | 指令控制版 | | CosyVoice2-0.5b | 第二代轻量版 | | FishSpeech | 高质量零样本语音克隆 | | IndexTTS | 工业级中文 TTS | | SparkTTS | 讯飞开源语音合成 | | GPT-SoVITS | 少样本声音克隆 |
→ aigcpanel 原生 TTS(7 项):CosyVoice-300M、CosyVoice-300M-Instruct、CosyVoice2-0.5b、FishSpeech、IndexTTS、SparkTTS、GPT-SoVITS。
"声音识别"表格:FunASR(阿里达摩院开源 ASR,支持带时间戳输出) "视频模型"表格:MuseTalk、LatentSync、Wav2Lip、Heygem
README 逐字:
- 本地模型一键导入、启动/停止、日志查看、参数配置
- 支持远程 API 模型接入
- 云端 AI 模型服务(无需本地显卡)(VIP)
代码侧(文件树来源 https://data.jsdelivr.com/v1/packages/gh/modstart-lib/aigcpanel@master)可见 TTS 走统一任务/工作流抽象:
/src/pages/Apps/LongTextTts/... (长文本转音频)
/src/pages/Apps/SubtitleTts/... (字幕转音频)
/src/pages/Sound/components/SoundGenerateForm.vue
/src/pages/Sound/components/SoundGenerateSelector.vue ← 模型选择器
/src/task/SoundGenerate.ts
→ 模型以"启动包/模型市场"形式可插拔,并支持远程 API 模型;但"接入任意第三方 TTS"的自定义插件文档:未查到。
| 项 | 值 |
|---|---|
| 仓库 | https://github.com/Soul-AILab/SoulX-FlashTalk |
| 星标 | {"label":"stars","message":"1.5k",...} |
| 许可证 | {"label":"license","message":"Apache-2.0",...} |
| 最后提交 | {"label":"last commit","message":"july",...} |
抓取地址:https://img.shields.io/github/{stars,license,last-commit}/Soul-AILab/SoulX-FlashTalk.json
来源:https://raw.githubusercontent.com/Soul-AILab/SoulX-FlashTalk/main/README.md(12367 B)
全文标题逐字:
SoulX-FlashTalk: Real-Time Infinite Streaming of Audio-Driven Avatars via Self-Correcting Bidirectional Distillation
chinese-wav2vec2-base)。|
SoulX-FlashTalk-14B| Our 14b model | 🤗 Huggingface | |chinese-wav2vec2-base| chinese-wav2vec2-base | 🤗 Huggingface |
chinese-wav2vec2-base(换音频编码器需改代码,文档未提供插件点)。Real-time inference speed can only be supported on 8xH800 or higher graphics cards
Requires more than 64G of VRAM. Use --cpu_offload to reduce VRAM usage to 40G.
generate_video.py 提供 --audio_encode_mode(逐字):help="stream: encode audio chunk before every generation; once: encode audio together"
2026.02.12 - We have released the SoulX-FlashHead, which is a streaming talking head project that achieves real-time performance on consumer GPUs (e.g., RTX 4090/5090).来源:https://cdn.jsdelivr.net/gh/Soul-AILab/SoulX-FlashTalk@main/flash_talk/configs/infer_params.yaml(逐字全文)
frame_num: 33
motion_frames_num: 5
tgt_fps: 25
sample_rate: 16000
sample_steps: 4
sample_shift: 5
color_correction_strength: 1.0
cached_audio_duration: 8
height: 720
width: 416
来源:https://cdn.jsdelivr.net/gh/Soul-AILab/SoulX-FlashTalk@main/generate_video.py(逐字)
human_speech_array_all, _ = librosa.load(args.audio_path, sr=infer_params['sample_rate'], mono=True)
human_speech_array_slice_len = slice_len * sample_rate // tgt_fps
human_speech_array_frame_num = frame_num * sample_rate // tgt_fps
default="examples/cantonese_16k.wav",
help="[meta file] The audio path to generate the video.")
librosa.load(..., sr=16000, mono=True))sample_rate // tgt_fps = 16000/25 = 640 个采样点(= 40 ms)cached_audio_duration: 8(缓存 8 秒音频)cantonese_16k.wav(16k)| 项 | 值 |
|---|---|
| 仓库 | https://github.com/anliyuan/Ultralight-Digital-Human |
| 星标 | {"label":"stars","message":"2.6k",...} |
| 许可证 | {"label":"license","message":"not specified",...} |
| 最后提交 | {"label":"last commit","message":"july",...} |
抓取地址:https://img.shields.io/github/{stars,license,last-commit}/anliyuan/Ultralight-Digital-Human.json
注:
raw.githubusercontent.com/.../main/README.md返回 404;默认分支为master,取自https://cdn.jsdelivr.net/gh/anliyuan/Ultralight-Digital-Human@master/README.md。
来源:https://cdn.jsdelivr.net/gh/anliyuan/Ultralight-Digital-Human@master/README.md
.npy,再喂模型(逐字):inference
Before run inference, you need to extract test audio feature(i will merge this step and inference step), run this 在推理之前,需要先提取测试音频的特征(之后会把这步和推理合并到一起去),运行(音频采样率需要是16000)
python data_utils/hubert.py --wav your_test_audio.wav # when using hubert or python data_utils/python wenet_infer.py your_test_audio.wav # when using wenetthen you get
your_test_audio_hu.npyoryour_test_audio_wenet.npypython inference.py --asr hubert --dataset ./your_data_dir/ --audio_feat your_test_audio_hu.npy --save_path xxx.mp4 --checkpoint your_trained_ckpt.pthTo merge the audio and the video, run
ffmpeg -i xxx.mp4 -i your_audio.wav -c:v libx264 -c:a aac result_test.mp4
上游 TTS 完全由用户自备(音频 → 特征 → 视频,再 ffmpeg 合音轨)。
wenet 或 hubert(--asr 参数),但这是音频特征提取器,不是 TTS。关于流式推理:
使用流式推理时,建议把静音的图片和对应的关键点放在单独的目录里,img_inference和lms_inference里。
因为一般用到流式推理的场景一般对实时性要求比较高,所以这里我只写了wenet作为音频编码器的情况(实测在2080这样的机器上多个并发时每帧音频处理+视频处理耗时10ms以内,需要将模型转为onnx)。
这个模型是支持流式推理的,但是代码还没有完善,之后我会提上来。
c++流式推理代码已开源到FeatherTalk🎉 … 它是 Ultralight Digital Human 的整理和升级版,重点优化了音频编码器、训练流程和移动端部署体验。
When you using wenet, you neet to ensure that your video frame rate is 20, and for hubert, your video frame rate should be 25. 如果你选择使用wenet的话,你必须保证你视频的帧率是20fps,如果选择hubert,视频帧率必须是25fps。
| 项 | 值 |
|---|---|
| 仓库 | https://github.com/Open-LLM-VTuber/Open-LLM-VTuber |
| 星标 | {"label":"stars","message":"14k",...} |
| 许可证 | {"label":"license","message":"not identifiable by github",...}(README 徽章指向 LICENSE,第三方页脚称 "the MIT license of this project") |
| 最后提交 | {"label":"last commit","message":"may",...} |
抓取地址:https://img.shields.io/github/{stars,license,last-commit}/Open-LLM-VTuber/Open-LLM-VTuber.json
来源:https://raw.githubusercontent.com/Open-LLM-VTuber/Open-LLM-VTuber/main/README.md(11536 B)
🧠 Extensive model support:
- 🤖 Large Language Models (LLM): Ollama, OpenAI (and any OpenAI-compatible API), Gemini, Claude, Mistral, DeepSeek, Zhipu AI, GGUF, LM Studio, vLLM, etc.
- 🎙️ Automatic Speech Recognition (ASR): sherpa-onnx, FunASR, Faster-Whisper, Whisper.cpp, Whisper, Groq Whisper, Azure ASR, etc.
- 🔊 Text-to-Speech (TTS): sherpa-onnx, pyttsx3, MeloTTS, Coqui-TTS, GPTSoVITS, Bark, CosyVoice, Edge TTS, Fish Audio, Azure TTS, etc.
→ 10 个已点名 TTS("etc." 表示还有更多):sherpa-onnx、pyttsx3、MeloTTS、Coqui-TTS、GPTSoVITS、Bark、CosyVoice、Edge TTS、Fish Audio、Azure TTS。
- 🔧 Highly customizable:
- ⚙️ Simple module configuration: Switch various functional modules through simple configuration file modifications, without delving into the code
- 🔌 Good extensibility: Modular design allows you to easily add your own LLM, ASR, TTS, and other module implementations, extending new features at any time
→ 配置文件切换 + 可自行新增 TTS 模块实现(本报告中可插拔性最好的一个)。
📢 v2.0 Development: We are focusing on Open-LLM-VTuber v2.0 — a complete rewrite of the codebase.)。weihaox/awesome-digital-humanhttps://cdn.jsdelivr.net/gh/weihaox/awesome-digital-human@main/README.md(39329 B)tts|sovits|cosyvoice|chattts|fish|index.?tts|edge|azure|paddle|mega|bark|xtts 全量 grep:未命中任何 TTS 后端(仅误命中 "MEGANE"(眼镜论文)与 "DiffCloth")。icemaple77/digital-humanhttps://cdn.jsdelivr.net/gh/icemaple77/digital-human@main/README.md(3404 B)四个模块(STT / LLM / TTS / 数字人)各自独立,换实现只改
.env一处,主程序不动。 ├─ TTS 文字→声音 (edge 免费 | GPT-SoVITS 自训音色) | 嘴 TTS |edge(免费无模型)/gptsovits(自训音色) |TTS_PROVIDER;gptsovits 配TTS_BASE_URL+参考音频 |
→ TTS 后端:edge(免费、无模型)/ gptsovits(自训音色,走 :9880 api_v2)。 通过 .env 的 TTS_PROVIDER 切换 —— 配置级可插拔(但仅两个内置实现)。
来源:https://blog.csdn.net/jacke121/article/details/157323875(标题《LiveTalking 部署笔记》)
该文在 LiveTalking 语境下给出「tts选型」表,逐字摘录(含首包延迟与显存):
| 算法名称 | 核心特点 | 克隆速度与效果 | 资源/硬件要求 | 一句话总结 |
|---|---|---|---|---|
| IndexTTS2 | 效果与资源的完美平衡,支持时长与情感精细控制 | 克隆效果好,可控制情感强度,首包延迟约1.2秒 | 8GB显存 即可流畅运行 | ✅ 非常适合,资源友好,效果出色 |
| FishTTS / FishSpeech | 速度之王,流式架构带来毫秒级响应 | 首帧延迟低于 500ms,10秒样本即可克隆 | 显存占用仅 3.2GB,轻量高效 | ✅ 非常适合,极速体验,资源占用极低 |
| CosyVoice | 全能选手,功能全面均衡,由阿里开源 | 克隆相似度达95%以上,流式首包延迟低于 150ms | 显存需求约 6GB | ⭐ 可考虑,表现均衡,但非克隆特长生 |
| XTTS | 多语言专家,一次训练支持17种语言 | 支持多语言零样本克隆,但社区实测克隆相似度略低于专注中文的模型 | 约 6-8GB | ⚠️ 除非你需要极强的多语言能力,否则不是首选 |
| SoVITS | 歌声转换,专为"AI翻唱"设计,非普通语音克隆 | 效果惊艳,但需长音频(10-30分钟)训练,推理显存占用大 | 显存占用大,6GB显存仅能处理30秒音频 | ❌ 不适合,场景是唱歌,且资源消耗大 |
| EdgeTTS | 便捷免费的云端服务,非本地克隆算法 | 无法克隆声音,只能使用微软预设的200多种音色 | 无需显卡,只需网络 | ❌ 不适合,不支持声音克隆 |
→ 与 LiveTalking 内置列表对照:IndexTTS2 / FishTTS / CosyVoice / XTTS / SoVITS(即 gpt-sovits) / EdgeTTS 全部都在 LiveTalking 的原生枚举里(见 §1.3)。也就是说社区推荐顺序 = IndexTTS2 ≈ FishSpeech > CosyVoice > XTTS > SoVITS/EdgeTTS。 该文另有一条 LiveTalking 实操原文(逐字):
默认的是edgetts,我选择采用gpt-sovits,edgetts相较于gpt-sovits生成的速度会快一点,但不过音色质量会相对弱一点,如果电脑配置支持,可以在本地部署一个gpt-sovits。
以及:
--REF_FILE:用于声音克隆的样例音频文件,--REF_TEXT:样例文本信息,--push_url:srs推流服务地址。
来源:https://juejin.cn/post/7649764223339495475 标题逐字:《调查研究-169 开源 TTS 模型横向对比:从"能发声"到"可部署的语音智能基础设施"(2026 版)》
摘要逐字:
场景:面向语音助手、陪伴机器人、短视频配音、有声书、数字人、客服播报等场景的 TTS 选型与工程落地。 结论:当前 TTS 已分化为四类——传统工程型(MeloTTS)、零样本克隆型(F5-TTS / Spark-TTS / CosyVoice / IndexTTS2)、LLM 化生成型(VoxCPM / Qwen3-TTS / Fish Audio S2 / CosyVoice)、生产服务型(Qwen TTS Realtime)。
其「版本矩阵」表(逐字,截取列)中的首包延迟与许可证:
| 模型 | 最新版本 | 发布时间 | 许可证 | 商业可用 | 首包延迟 |
|---|---|---|---|---|---|
| VoxCPM2 | VoxCPM2 | 2026-04 | Apache-2.0 | ✅ | RTF ≈ 0.13 |
| Qwen3-TTS | 1.7B / 0.6B | 2026-01-22 | Apache-2.0 | ✅ | 端到端 97ms(首字符即出) |
| Qwen TTS Realtime | 云服务 | 持续迭代 | 阿里云商业 | ✅(按量计费) | < 200ms(云端 SLA) |
| CosyVoice 2 | CosyVoice2-0.5B | 2024-12 | Apache-2.0 | ✅ | 首包 150ms |
| F5-TTS | 1.x | 2024 | MIT | ✅ | 取决于推理后端 |
| IndexTTS2 | 2.0 | 2025-09 | Apache-2.0 | ✅ | 取决于部署 |
| Fish Audio S2 / S2-Pro | S2 / S2-Pro | 2026-03-11 | S2 相对开放 / S2-Pro Research License | ⚠️ S2-Pro 商用需单独授权 | < 150ms(官方指标) |
| Spark-TTS | 0.5B | 2025 | Apache-2.0 | ✅ | — |
| MeloTTS | 持续维护 | 2024-12 末次主提交 | MIT | ✅ | CPU 实时 |
该文对数字人场景的直接结论(逐字):
第二类是零样本声音克隆型 TTS。代表包括 F5-TTS、Spark-TTS、CosyVoice、IndexTTS2 等。… 这类模型很适合短视频配音、有声书、个人声音助手、数字人和角色语音。
IndexTTS2 … 它的另一个关键能力是时长控制。视频配音和口型同步经常要求一句话必须在固定时间内说完。 … IndexTTS2 的 duration control 对影视、动画、虚拟人、字幕对齐非常有意义。
总体判断:IndexTTS2 适合做情绪配音、角色语音、影视/视频配音、有声书和数字人内容生成。它的上限很高,但商业集成要谨慎。
CosyVoice … 总体判断:CosyVoice 是最适合作为"本地中文 TTS 工程底座"的模型之一。 它未必在每个单项上都第一,但综合能力强,适合认真投入。
总体判断:F5-TTS 适合做轻量自部署、声音克隆、内容生产和实验平台。它是"好用的生成工具",但未必是"实时语音基础设施"的最优解。
Fish Audio S2 Pro 模型卡显示 Research License,非商业免费,商业使用需要单独授权。
补充说明(诚实标注):该文的许可证/延迟数据来自第三方博客转述,本次调研未逐一回溯其原始模型卡;引用时请以该博文为来源。
| 社区 2026 推荐 | LiveTalking | OpenAvatarChat | aigcpanel | 备注 |
|---|---|---|---|---|
| IndexTTS2 | ✅ indextts2 | ❌ | ✅ IndexTTS | LiveTalking 文档给"延时4s / 17G"(OmniTTS 路径),CSDN 实测"首包约1.2s / 8G" |
| FishSpeech | ✅ fishtts | ❌ | ✅ FishSpeech | 三处一致 |
| CosyVoice | ✅ cosyvoice | ✅ 本地 + 百炼 | ✅ 3 个版本 | 覆盖最广 |
| GPT-SoVITS | ✅ gpt-sovits | ❌ | ✅ | — |
| Edge TTS | ✅ edgetts(默认) | ✅ | ❌(不在声音合成表) | LiveTalking / OAC 的零成本起步项 |
| Qwen3-TTS / VoxCPM2 | ✅(经 OmniTTS/vLLM-Omni) | ⚠️ 仅 Qwen-Omni 多模态路径 | ❌ | LiveTalking 通过统一 OmniTTS 网关接新模型 |
| SparkTTS | ❌ | ❌ | ✅ | 仅 aigcpanel |
| F5-TTS / MeloTTS / Bark / XTTS | XTTS ✅(xtts,未列入 config 注释) | ❌ | ❌ | Open-LLM-VTuber 覆盖 F5?/MeloTTS/Bark |
共性结论(2026-09):
CosyVoice + Edge TTS 几乎是所有框架的最小公倍数(LiveTalking / OpenAvatarChat / Linly-Talker / aigcpanel / Open-LLM-VTuber 全部支持)。GPT-SoVITS / FishSpeech / IndexTTS2 是第二梯队,在 LiveTalking 与 aigcpanel 里都有原生位。sample_rate: 16000;Ultralight 明确"音频采样率需要是16000")。sample_rate = 16000、chunk = 16000//(fps*2) = 320 samples = 20ms)。https://img.shields.io/github/stars/lipku/LiveTalking.jsonhttps://img.shields.io/github/license/lipku/LiveTalking.jsonhttps://img.shields.io/github/last-commit/lipku/LiveTalking.jsonhttps://img.shields.io/github/stars/HumanAIGC/OpenAvatarChat.json(返回 repo not found → 用于确认 owner 变更)https://img.shields.io/github/stars/HumanAIGC-Engineering/OpenAvatarChat.jsonhttps://img.shields.io/github/license/HumanAIGC-Engineering/OpenAvatarChat.jsonhttps://img.shields.io/github/last-commit/HumanAIGC-Engineering/OpenAvatarChat.jsonhttps://img.shields.io/github/stars/xszyou/Fay.jsonhttps://img.shields.io/github/license/xszyou/Fay.jsonhttps://img.shields.io/github/last-commit/xszyou/Fay.jsonhttps://img.shields.io/github/stars/Korvo-AI/Linly-Talker.json(repo not found → owner 变更)https://img.shields.io/github/stars/Kedreamix/Linly-Talker.jsonhttps://img.shields.io/github/license/Kedreamix/Linly-Talker.jsonhttps://img.shields.io/github/last-commit/Kedreamix/Linly-Talker.jsonhttps://img.shields.io/github/stars/duixcom/Duix.Heygem.jsonhttps://img.shields.io/github/license/duixcom/Duix.Heygem.jsonhttps://img.shields.io/github/last-commit/duixcom/Duix.Heygem.jsonhttps://img.shields.io/github/stars/duixcom/Duix.Avatar.jsonhttps://img.shields.io/github/license/duixcom/Duix.Avatar.jsonhttps://img.shields.io/github/last-commit/duixcom/Duix.Avatar.jsonhttps://img.shields.io/github/stars/modstart-lib/aigcpanel.jsonhttps://img.shields.io/github/license/modstart-lib/aigcpanel.jsonhttps://img.shields.io/github/last-commit/modstart-lib/aigcpanel.jsonhttps://img.shields.io/github/stars/Soul-AILab/SoulX-FlashTalk.jsonhttps://img.shields.io/github/license/Soul-AILab/SoulX-FlashTalk.jsonhttps://img.shields.io/github/last-commit/Soul-AILab/SoulX-FlashTalk.jsonhttps://img.shields.io/github/stars/anliyuan/Ultralight-Digital-Human.jsonhttps://img.shields.io/github/license/anliyuan/Ultralight-Digital-Human.jsonhttps://img.shields.io/github/last-commit/anliyuan/Ultralight-Digital-Human.jsonhttps://img.shields.io/github/stars/Open-LLM-VTuber/Open-LLM-VTuber.jsonhttps://img.shields.io/github/license/Open-LLM-VTuber/Open-LLM-VTuber.jsonhttps://img.shields.io/github/last-commit/Open-LLM-VTuber/Open-LLM-VTuber.jsonhttps://raw.githubusercontent.com/lipku/LiveTalking/main/README.mdhttps://raw.githubusercontent.com/lipku/LiveTalking/main/README-EN.mdhttps://raw.githubusercontent.com/HumanAIGC-Engineering/OpenAvatarChat/main/README.mdhttps://raw.githubusercontent.com/HumanAIGC-Engineering/OpenAvatarChat/main/readme_en.mdhttps://raw.githubusercontent.com/HumanAIGC/OpenAvatarChat/main/README.md(失败:仓库不存在)https://raw.githubusercontent.com/xszyou/Fay/master/README.mdhttps://raw.githubusercontent.com/xszyou/Fay/master/system.conf.bakhttps://raw.githubusercontent.com/xszyou/Fay/master/tts/ali_tss.pyhttps://raw.githubusercontent.com/xszyou/Fay/master/tts/ms_tts_sdk.pyhttps://raw.githubusercontent.com/xszyou/Fay/master/tts/tts_voice.pyhttps://raw.githubusercontent.com/Kedreamix/Linly-Talker/main/README.mdhttps://raw.githubusercontent.com/Korvo-AI/Linly-Talker/main/README.md(失败:仓库不存在)https://raw.githubusercontent.com/duixcom/Duix.Heygem/main/README.mdhttps://raw.githubusercontent.com/modstart-lib/aigcpanel/main/README.mdhttps://raw.githubusercontent.com/Soul-AILab/SoulX-FlashTalk/main/README.mdhttps://raw.githubusercontent.com/Open-LLM-VTuber/Open-LLM-VTuber/main/README.mdhttps://raw.githubusercontent.com/anliyuan/Ultralight-Digital-Human/main/README.md(404 → 默认分支为 master)https://data.jsdelivr.com/v1/packages/gh/lipku/LiveTalking@mainhttps://data.jsdelivr.com/v1/packages/gh/HumanAIGC-Engineering/OpenAvatarChat@mainhttps://data.jsdelivr.com/v1/packages/gh/Soul-AILab/SoulX-FlashTalk@mainhttps://data.jsdelivr.com/v1/packages/gh/modstart-lib/aigcpanel@masterhttps://data.jsdelivr.com/v1/packages/gh/Kedreamix/Linly-Talker@main(403)https://data.jsdelivr.com/v1/packages/gh/xszyou/Fay@master(403)https://data.jsdelivr.com/v1/packages/gh/duixcom/Duix.Heygem@main(403)https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/config.yamlhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/config.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/registry.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/requirements.txthttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/__init__.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/base_tts.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/edge.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/azure.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/doubao.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/omnitts.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/qwentts.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/sovits.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/tencent.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/xtts.pyhttps://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/cosyvoice.py(404)https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/fish.py(404)https://cdn.jsdelivr.net/gh/lipku/LiveTalking@main/tts/indextts2.py(404)https://cdn.jsdelivr.net/gh/HumanAIGC-Engineering/OpenAvatarChat@main/src/handlers/tts/cosyvoice/tts_handler_cosyvoice.pyhttps://cdn.jsdelivr.net/gh/HumanAIGC-Engineering/OpenAvatarChat@main/src/handlers/tts/edgetts/tts_handler_edgetts.pyhttps://cdn.jsdelivr.net/gh/HumanAIGC-Engineering/OpenAvatarChat@main/config/chat_with_openai_compatible_edge_tts.yamlhttps://cdn.jsdelivr.net/gh/xszyou/Fay@master/tts/ali_tss.pyhttps://cdn.jsdelivr.net/gh/xszyou/Fay@master/tts/ms_tts_sdk.pyhttps://cdn.jsdelivr.net/gh/xszyou/Fay@master/tts/tts_voice.pyhttps://cdn.jsdelivr.net/gh/Kedreamix/Linly-Talker@main/TTS/README.mdhttps://cdn.jsdelivr.net/gh/anliyuan/Ultralight-Digital-Human@master/README.mdhttps://cdn.jsdelivr.net/gh/Soul-AILab/SoulX-FlashTalk@main/flash_talk/configs/infer_params.yamlhttps://cdn.jsdelivr.net/gh/Soul-AILab/SoulX-FlashTalk@main/generate_video.pyhttps://cdn.jsdelivr.net/gh/weihaox/awesome-digital-human@main/README.mdhttps://cdn.jsdelivr.net/gh/icemaple77/digital-human@main/README.mdhttps://doc.livetalking.ai/docs/tts/https://doc.livetalking.ai/docs/tts/gptsovits/https://doc.livetalking.ai/docs/tts/fishspeech/https://doc.livetalking.ai/docs/tts/cosyvoice/https://doc.livetalking.ai/docs/tts/omnitts/https://doc.livetalking.ai/docs/faq/https://humanaigc-engineering.github.io/OpenAvatarChat/reference/preset-modeshttps://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/bailian-cosyvoice.htmlhttps://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/cosyvoice-local.htmlhttps://humanaigc-engineering.github.io/OpenAvatarChat/reference/handlers/tts/edge-tts.htmlhttps://deepwiki.com/xszyou/Fay/4.2-text-to-speech-(tts)https://deepwiki.com/xszyou/Fay/10.1-local-setup-and-configurationhttps://deepwiki.com/xszyou/Fayhttps://deepwiki.com/xszyou/Fay/4-inputoutput-systemshttps://juejin.cn/post/7649764223339495475 —《开源 TTS 模型横向对比:从"能发声"到"可部署的语音智能基础设施"(2026 版)》,2026-06-12https://blog.csdn.net/jacke121/article/details/157323875 —《LiveTalking 部署笔记》,2026-01-24 首发 / 2026-03-30 更新| 框架 | 未查到内容 |
|---|---|
| LiveTalking | TTS 级 TTFB 硬阈值(FAQ 未涉及 TTS);cosyvoice.py / fish.py / indextts2.py 的注册名原文(CDN 404) |
| OpenAvatarChat | TTS 级延迟数字(只有系统级"平均响应时间 2.2 秒");TTS chunk 大小约束 |
| Fay | TTS 级 TTFB / ms 指标;火山引擎(volcano)输出的采样率(README 与源码未在本轮取到该实现细节) |
| Linly-Talker | TTS 输出采样率 / chunk 约束(README 与 TTS/README.md 均未声明) |
| Duix.Avatar | 采样率要求;TTFB(离线任务式);TTS 替换机制文档 |
| aigcpanel | 任何延迟 / FPS / 流式指标;采样率 / chunk 约束;自定义 TTS 插件文档 |
| SoulX-FlashTalk | TTS 相关项不适用(无 TTS) |
| Ultralight-Digital-Human | TTS 相关项不适用(无 TTS);每帧音频 chunk 大小 |
| Open-LLM-VTuber | TTS 级延迟数字;采样率约束 |
| awesome 清单 | weihaox/awesome-digital-human 不列任何 TTS 后端 |
报告结束。所有引用均来自上列 URL 的实际抓取内容;凡 jsDelivr 与 raw 存在分歧处(Linly-Talker / Ultralight / OpenAvatarChat 的文件树 vs CDN),已在对应章节显式标注。