| 目标 | 状态 | 说明 |
|---|---|---|
github.com HTML 页面 | 不通/极慢(curl 60s 超时;偶发 200) | 因此无法用 GitHub 网页正文做证据 |
raw.githubusercontent.com | 通(偶发慢/失败,需重试) | README 正文来源 |
cdn.jsdelivr.net(GitHub 镜像 CDN) | 稳定通 | README/LICENSE 正文主要来源 |
img.shields.io GitHub badge JSON | 稳定通 | star / license / last-commit 来源 |
api.github.com(search/issues) | 通,未认证限流 ~10 req/min | issue 检索来源 |
huggingface.co | 屏蔽 | 改用 hf-mirror.com |
hf-mirror.com(含 /api/models) | 通 | 模型卡正文、safetensors 参数量、license tag |
export.arxiv.org/api/query | 通 | 论文标题/摘要/发布日期的权威验证源 |
arxiv.org/abs/* | 通(200) | 人类可读链接 |
本文对每条 URL 标注了核实方式,只有三种:
[正文已抓取] — 我用 curl 真正下载了该 URL 的正文(HTTP 200 且内容非平凡),并据此写了结论。[shields-verified] — 通过 img.shields.io/github/{stars,license,last-commit}/<owner>/<repo>.json 返回了真实数据(而非 "repo not found"),据此确认仓库存在及其 star/license/最后提交。这是本机在 github.com 网页不可达时唯一能确认仓库真实性的手段。[arxiv-api-verified] — 通过 https://export.arxiv.org/api/query?id_list=<id> 返回的 XML 中 <title> 与 <published> 与本文引用一致。本文逐一标注每个数字的来源类型:
LICENSE 是 Apache-2.0,但 HuggingFace 权重模型卡的 license 是 cc-by-nc-sa-4.0(禁止商用)。代码可商用、权重不可商用 → 整体判定 不可商用。见 §5。boson-ai/higgs-audio 仓库 README 现已改为宣传 Higgs Audio v3,其许可为 Boson Higgs Audio v3 Research and Non-Commercial License(明确禁止商用);v2 文档被挪到 README_V2.md(Apache-2.0 代码 + HF license: other)。且 Higgs Audio v2/v3 均未查到 arXiv 论文(只有博客)。见 §6。VibeVoice-TTS-1.5B 在官方表里状态为 Disabled;当前活跃的是 VibeVoice-Realtime-0.5B(流式 TTS)和 VibeVoice-ASR 系列。且官方 README 明确写"不建议用于商业或真实场景"。见 §7。结论速览:Apache-2.0 全开源、2B、48kHz、~8GB 显存、RTF ~0.30(4090),可商用,是"数字人 + 中文 + 实时"最稳的开源底座之一。但 sm_120(RTX 50 系 Blackwell)有若干未关闭的 issue。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/OpenBMB/VoxCPM | [shields-verified] |
| Star 数 | 38k | [shields-verified] https://img.shields.io/github/stars/OpenBMB/VoxCPM.json → {"message":"38k"} |
| License(仓库) | Apache-2.0 | [shields-verified] + [正文已抓取] https://cdn.jsdelivr.net/gh/OpenBMB/VoxCPM@main/LICENSE(11303 B,Apache License 2.0 全文) |
| Last commit | september(2026-09) | [shields-verified] https://img.shields.io/github/last-commit/OpenBMB/VoxCPM.json → {"message":"september"} |
| HuggingFace | https://huggingface.co/openbmb/VoxCPM2 | [正文已抓取](镜像)https://hf-mirror.com/openbmb/VoxCPM2/raw/main/README.md(200,7939 B) |
| HF license tag | apache-2.0;gated: False | [正文已抓取] https://hf-mirror.com/api/models/openbmb/VoxCPM2?blobs=false → license: ['license:apache-2.0'] |
✅ 补充确认(2026-09-22 追加):VoxCPM2 的 HF 侧许可证
提问背景:抓
openbmb/VoxCPM返回 401,无法独立确认 HF 侧许可。原因是 repo 名不对。
尝试的 repo 名 结果 结论 https://hf-mirror.com/OpenBMB/VoxCPM/raw/main/README.mdHTTP 401 ❌ 旧名/错名(大写 owner + 无 "2")→ 这就是 401 的来源 https://hf-mirror.com/openbmb/VoxCPM2/raw/main/README.mdHTTP 200 ✅ 正确的 HF repo 是 openbmb/VoxCPM2(owner 全小写,模型名带2)HF 模型卡 frontmatter 第 33 行原文:
license: apache-2.0([正文已抓取],sed -n '1,40p'实测) HF API tags 原文:license: ['license:apache-2.0'],gated: False(无需门禁,可直接下载)→ 结论:VoxCPM2 的 HF 侧许可证确实为 Apache-2.0,与 GitHub README 声称的 "free for commercial use" 一致。HF 与 GitHub 两侧无冲突。可商用。
另外注意:
https://hf-mirror.com/OpenBMB/VoxCPM2/...(仅 owner 大写)也会返回 200(HF owner 名大小写不敏感),但你试的OpenBMB/VoxCPM缺了2且是旧名,所以 401。请统一使用openbmb/VoxCPM2。
| arXiv(VoxCPM2) | https://arxiv.org/abs/2606.06928 = VoxCPM2 Technical Report,2026-06-05 | [arxiv-api-verified] | | arXiv(VoxCPM 初代) | https://arxiv.org/abs/2509.24650 = VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning,2025-09-29 | [arxiv-api-verified] |
| 项 | 值 | 来源 |
|---|---|---|
| 骨干参数量 | 2B(README 明示) | 【厂商自报】[正文已抓取] README "Models & Versions" 表:Backbone Parameters = 2B |
| HF 权重实际张量总量 | 2,290,004,544 ≈ 2.29B(BF16) | [正文已抓取] https://hf-mirror.com/api/models/openbmb/VoxCPM2?blobs=false → safetensors.parameters.BF16 |
| 采样率 | 48 kHz 输出(AudioVAE V2,16 kHz 输入 → 48 kHz 输出,内置超分) | 【厂商自报】HF 模型卡 |
| 语言 | 30 语言 + 9 种中文方言(四川话/粤语/吴语/东北话/河南话/陕西话/山东话/天津话/闽南话) | 【厂商自报】HF 模型卡 |
| LM token rate | 6.25 Hz | 【厂商自报】HF 模型卡 |
| dtype | bfloat16 | 【厂商自报】HF 模型卡 |
| VRAM | ~8 GB | 【厂商自报】HF 模型卡 Model Details 表 + GitHub README 对比表 |
Seed-TTS-eval(【厂商自报】,VoxCPM2 GitHub README 对比表):
| 指标 | VoxCPM2 (2B) | 参照 |
|---|---|---|
| test-ZH CER ↓ | 0.97 % | Qwen3-TTS 1.7B 1.22 / CosyVoice3-1.5B 1.12 / Seed-TTS 1.12 / MiniMax-Speech 0.83 |
| test-ZH SIM ↑ | 79.5 % | Seed-TTS 79.6 / MiniMax-Speech 78.3 / CosyVoice3-1.5B 78.1 |
| test-EN WER ↓ / SIM ↑ | 1.84 / 75.3 | — |
| test-Hard CER ↓ / SIM ↑ | 8.13 / 75.3 | Seed-TTS 7.59 / CosyVoice3-1.5B 5.83 |
CV3-eval 多语种 CER/WER(【厂商自报】):zh 3.65、en 5.00、hard-zh 8.55、hard-en 8.48、ja 5.96、ko 5.69、de 4.77、es 3.80、fr 9.85、it 4.25、ru 5.21。
MiniMax-Multilingual-Test 中文 WER(【厂商自报】):VoxCPM2 1.136;Qwen3-TTS 0.928;FishAudio S2 0.730;MiniMax 2.252;ElevenLabs 16.026。
内部 30 语种评测(论文摘要):平均 WER 1.68 %(【厂商自报】)。
UTMOS / 自然度:未查到。
| 项 | 值 | 来源 |
|---|---|---|
| 流式支持 | 支持 — Python API model.generate_streaming(text=...) 逐块 yield;CLI --stream | 【厂商自报】[正文已抓取] GitHub README |
| RTF(RTX 4090,标准 PyTorch) | ~0.30 | 【厂商自报】 |
| RTF(RTX 4090,Nano-VLLM) | ~0.13 | 【厂商自报】 |
| RTF(Apple M4 Pro / Metal,Q8_0,llama.cpp-omni) | ~1.76 | 【厂商自报】 |
| TTFB / 首包延迟(ms) | 未查到(README 只给 RTF,不给首包 ms) | — |
| 版本对比 | VoxCPM2 RTF ~0.30 / VRAM ~8GB;VoxCPM1.5 RTF ~0.15 / ~6GB;VoxCPM-0.5B RTF ~0.17 / ~5GB | 【厂商自报】README 对比表 |
环境要求【厂商自报】:Python ≥ 3.10 (<3.13)、PyTorch ≥ 2.5.0、CUDA ≥ 12.0。
GitHub issue 检索(https://api.github.com/search/issues?q=repo:OpenBMB/VoxCPM+...):
| Issue | 标题 | 状态 | 对数字人的影响 |
|---|---|---|---|
| #250 | feat: add Docker support for NVIDIA RTX 5000 series (Blackwell / sm_120) | open | sm_120 官方 Docker 支持仍缺失 |
| #326 | LoRA fine-tuning + nano-vllm on Blackwell (sm_120): progressive audio quality degradation due to CUDA graph + object leak | closed | sm_120 上 CUDA graph 导致音质渐进劣化(微调场景) |
| #269 | CUDA allocator / cudagraph race when VoxCPM2 runs in ≥2 concurrent subprocess workers on the same GPU | closed | 多并发 worker 有 allocator 竞态 |
| #282 | Low Inference speed under RTX pro 6000 blackwell compared to RTX 4090 | closed | Blackwell 上反而比 4090 慢(用户报告) |
| #102 | Segfault on RTX 4090 during model warm-up (scaled_dot_product_attention) | closed | warm-up 段错误 |
是。 依据:原生 streaming API;RTF ~0.30(单张 4090 即 <1,可实时);~8GB 显存;官方提供 vLLM-Omni OpenAI 兼容 /v1/audio/speech 服务化路径(vllm serve openbmb/VoxCPM2 --omni);Apache-2.0 明确可商用(README 原文:"Fully Open-Source & Commercial-Ready — Weights and code released under the Apache-2.0 license, free for commercial use")。 风险:sm_120 上无官方 Docker + 多个 Blackwell 相关 issue(#250/#282/#326),RTX 50 系需自行验证。
结论速览:Apache-2.0、1.7B/0.6B、官方标注端到端合成延迟低至 97 ms、中文 Seed-TTS CER 0.77(12Hz-1.7B-Base),可商用,且 本次检索未发现任何 sm_120/Blackwell 问题——对 RTX 50 系最友好。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/QwenLM/Qwen3-TTS | [shields-verified] |
| Star 数 | 13k | [shields-verified] .../stars/QwenLM/Qwen3-TTS.json → {"message":"13k"} |
| License | Apache-2.0 | [正文已抓取] https://cdn.jsdelivr.net/gh/QwenLM/Qwen3-TTS@main/LICENSE(11343 B,Apache 2.0 全文) |
| Last commit | march | [shields-verified] |
| HuggingFace | https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice(同系列另有 -Base、-VoiceDesign、0.6B-*、Qwen/Qwen3-TTS-Tokenizer-12Hz) | [正文已抓取] https://hf-mirror.com/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice/raw/main/README.md(200,57846 B) |
| arXiv 2601.15621 核验 | ✅ 确认为真:http://arxiv.org/abs/2601.15621v1,标题 "Qwen3-TTS Technical Report",发布 2026-01-22 | [arxiv-api-verified] via https://export.arxiv.org/api/query?id_list=2601.15621 |
任务书疑问解答:IndexTTS 引用的 arXiv 2601.15621 确为 Qwen3-TTS 技术报告,不是张冠李戴。Qwen3-TTS GitHub README 的 Paper 徽章亦直接指向
https://arxiv.org/abs/2601.15621。
| 项 | 值 | 来源 |
|---|---|---|
| 已发布规模 | 1.7B / 0.6B 两档(另有 25Hz 版仅出现在论文评测表中,README 说"其他模型将陆续发布") | 【厂商自报】README 发布表 |
Qwen3-TTS-12Hz-1.7B-Base 实际张量数 | 1,928,677,440 ≈ 1.93B(BF16) | [正文已抓取] https://hf-mirror.com/api/models/Qwen/Qwen3-TTS-12Hz-1.7B-Base?blobs=false |
Qwen3-TTS-12Hz-1.7B-CustomVoice 实际张量数 | 1,916,676,352 ≈ 1.92B(BF16) | [正文已抓取] https://hf-mirror.com/api/models/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice?blobs=false |
Qwen3-TTS-12Hz-0.6B-Base 实际张量数 | 914,643,008 ≈ 0.91B(BF16) | [正文已抓取] https://hf-mirror.com/api/models/Qwen/Qwen3-TTS-12Hz-0.6B-Base?blobs=false |
Qwen3-TTS-Tokenizer-12Hz 实际张量数 | 170,557,441 ≈ 0.17B | [正文已抓取] https://hf-mirror.com/api/models/Qwen/Qwen3-TTS-Tokenizer-12Hz?blobs=false |
✅ 补充确认(2026-09-22 追加):Qwen3-TTS 的 HF repo 与许可证(应要求优先补)
HF 官方仓库集合:
https://huggingface.co/collections/Qwen/qwen3-tts(README 徽章指向)已逐一用 HF API 核实(
https://hf-mirror.com/api/models/<id>?blobs=false):
HF repo ID license tag gated 精确参数量 Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoicelicense:apache-2.0False 1,916,676,352 Qwen/Qwen3-TTS-12Hz-1.7B-Baselicense:apache-2.0False 1,928,677,440 Qwen/Qwen3-TTS-12Hz-0.6B-Baselicense:apache-2.0False 914,643,008 Qwen/Qwen3-TTS-Tokenizer-12Hzlicense:apache-2.0False 170,557,441 完整清单(README 发布表):
Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign、...-1.7B-CustomVoice、...-1.7B-Base、...-0.6B-CustomVoice、...-0.6B-Base、Qwen/Qwen3-TTS-Tokenizer-12Hz→ 结论:Qwen3-TTS 全部 HF repo 均为
apache-2.0且无门禁(gated: False),GitHub 仓库 LICENSE 亦为 Apache-2.0 全文。两侧一致,可商用,可直接下载无需授权。
| 语音 tokenizer | Qwen3-TTS-Tokenizer-12Hz(自研,12 Hz,离散多码本 LM 架构) | 【厂商自报】 | | 语言 | 10 种(中/英/日/韩/德/法/俄/葡/西/意)+ 多方言音色 | 【厂商自报】 | | VRAM | 未查到(README 只说"recommend using FlashAttention 2 to reduce GPU memory usage",未给 GB 数) | — |
Zero-shot on Seed-TTS test set(【厂商自报】README 评测表,"Seeds"行):
| 模型 | test-zh WER(即CER) ↓ | test-en WER ↓ |
|---|---|---|
| Seed-TTS | 1.12 | 2.25 |
| CosyVoice 3 | 0.71 | 1.45 |
| MiniMax-Speech | 0.83 | 1.65 |
| FireRedTTS 2 | 1.14 | 1.95 |
| Qwen3-TTS-12Hz-0.6B-Base | 0.92 | 1.32 |
| Qwen3-TTS-12Hz-1.7B-Base | 0.77 | 1.24 |
TTS 多语种测试集(【厂商自报】):Chinese 内容一致性 Qwen3-TTS-12Hz-1.7B-Base 0.928(Qwen3-TTS-25Hz-1.7B 为 0.777);Chinese Speaker Similarity 12Hz-0.6B 0.811、12Hz-1.7B 0.799(MiniMax 0.780、ElevenLabs 0.677)。
第三方交叉验证(【第三方表】VoxCPM2 README / MOSS-TTS README):两家的表都把 Qwen3-TTS-1.7B 记为 test-zh CER 1.22 / SIM 77.0(VoxCPM2 表);MOSS 表记 Qwen3-TTS-1.7B:EN WER 1.5、EN SIM 71.45、ZH CER 1.33、ZH SIM 76.72;Qwen3-TTS-0.6B:ZH CER 1.23、ZH SIM 76.4。 👉 注意口径差异:Qwen 自报 0.77 vs 第三方测 1.22~1.33。差异很可能来自是否使用 language="auto"、max_new_tokens、以及是否用官方 test-zh 参考音色。选型请以第三方口径为准做保守估计。
UTMOS / 自然度:未查到(Qwen3-TTS README 未给 UTMOS)。
| 项 | 值 | 来源 |
|---|---|---|
| 流式支持 | 支持 — "Dual-Track hybrid streaming generation architecture,单个模型同时支持流式与非流式" | 【厂商自报】README |
| 端到端合成延迟 | 低至 97 ms — 原文:"It can output the first audio packet immediately after a single character is input, with end-to-end synthesis latency as low as 97ms" | 【厂商自报】README(最强卖点,也是本次调研中最低的官方数字) |
| RTF | 未查到(README 未给 RTF) | — |
环境要求【厂商自报】:Python 3.12 fresh conda env;pip install -U qwen-tts;推荐 FlashAttention 2(pip install -U flash-attn --no-build-isolation;<96GB RAM 用 MAX_JOBS=4)。
GitHub issue 检索:
| Issue | 标题 | 状态 |
|---|---|---|
| #372 | finetuning/sft_12hz.py hard-codes attn_implementation="flash_attention_2" — official fine-tuning fails out-of-the-box | open |
| #373 | fix(finetuning): make attention implementation configurable | open |
| #350 | Multi-concurrency inference not supported | open |
| #369/#370 | 25Hz tokenizer: qkv_attention_manual never masks padded keys | open |
| #124 | macOS/M-Series Support for Qwen3-TTS | open |
| #345 | Add MLX / Apple Silicon (MPS) backend support | open |
| #179 | Finetuning Base results in progressively faster speech with every epoch | open |
⚠️ 关键坑:#372 —— 官方微调脚本硬编码 FA2,在没装 flash-attn 或硬件不支持 FA2(含部分 Blackwell 环境)时开箱即失败。
✅ sm_120 / Blackwell 专项检索结果:repo:QwenLM/Qwen3-TTS+sm_120 → total_count = 0(无相关问题报告)。这是本文所有模型中唯一 sm_120 零 issue 的。
非常适合。 依据:97 ms 官方端到端延迟(明确为"real-time interactive scenarios"设计);原生流式;Apache-2.0 可商用;无 sm_120 issue;提供 DashScope API 与 vLLM 路径。 风险:VRAM 未公布需实测;微调链路有硬编码 FA2 的坑(若不微调则不受影响)。
结论速览:Apache-2.0、1.5B、首包延迟低至 140 ms(L20)、bf16 后 9GB 显存、主打多人对话/播客,可商用,非常适合"多角色数字人对话/直播"。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/FireRedTeam/FireRedTTS2 | [shields-verified] |
| Star 数 | 1.4k | [shields-verified] |
| License | Apache-2.0 | [shields-verified] + [正文已抓取] https://cdn.jsdelivr.net/gh/FireRedTeam/FireRedTTS2@main/LICENSE(11357 B,Apache 2.0 全文) |
| Last commit | october 2025 | [shields-verified] |
| HuggingFace | https://huggingface.co/FireRedTeam/FireRedTTS2 | [正文已抓取] https://hf-mirror.com/FireRedTeam/FireRedTTS2/raw/main/README.md(200,8883 B);HF license tag = apache-2.0 |
| arXiv | https://arxiv.org/abs/2509.02020 = FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot,2025-09-02(v2) | [arxiv-api-verified] |
| 前代 | https://github.com/FireRedTeam/FireRedTTS(925 stars,MPL-2.0,last commit may 2025) | [shields-verified] |
| 项 | 值 | 来源 |
|---|---|---|
| 参数量 | 1.5B | 【第三方表】VoxCPM2 README 与 MOSS-TTS README 对比表均记 FireRedTTS-2 = 1.5B(两家独立一致) |
| HF 仓库 checkpoint 文件 | llm_pretrain.pt 8.27 GB、llm_posttrain.pt 8.27 GB、codec.pt 4.30 GB(合计 20.84 GB) | [正文已抓取] https://hf-mirror.com/api/models/FireRedTeam/FireRedTTS2/tree/main |
| VRAM | bf16 推理 9 GB(fp32 时 14 GB);原文 2025/09/28:"Supports bf16 inference, reducing VRAM usage from 14GB to 9GB and enabling consumer-grade GPU deployment." | 【厂商自报】README News |
| 采样 tokenizer | 12.5 Hz streaming speech tokenizer(自研) | 【厂商自报】 |
| 语言 | 英、中、日、韩、法、德、俄;支持跨语种/语码混合零样本克隆 | 【厂商自报】 |
| 能力上限 | 单次 3 分钟对话 / 4 个说话人(可通过扩数据线性扩展) | 【厂商自报】README Highlights |
Seed-TTS-eval:
UTMOS / 自然度:未查到(只有主观 MOS 的定性描述)。
| 项 | 值 | 来源 |
|---|---|---|
| 流式支持 | 支持 — 12.5 Hz 流式 tokenizer + text-speech 交织双 transformer,逐句生成;2025/10/11 起支持 streaming dialogue generation | 【厂商自报】README |
| 首包延迟(first-packet latency) | 低至 140 ms(NVIDIA L20 GPU) | 【厂商自报】README Highlights 原文 |
| RTF | 未查到(未给 RTF 数字) | — |
环境要求【厂商自报】:conda python==3.11;torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1 --index-url https://download.pytorch.org/whl/cu126;pip install -e . && pip install -r requirements.txt。
repo:FireRedTeam/FireRedTTS2 检索无相关结果)。[ ] 未勾选)。是,且是多说话人场景的首选之一。 依据:140 ms 首包(比 Qwen3-TTS 的 97 ms 略高但同量级);原生流式对话;9 GB 显存可上消费级卡;支持 4 说话人 3 分钟长对话;Apache-2.0 可商用;提供 Gradio UI 与微调教程。 风险:torch/cu126 版本锁定;中文 CER 1.14 弱于 Qwen3-TTS/VoxCPM2(单句质量不是它的强项,它的强项是对话与流式)。
结论速览:Apache-2.0,但只开源了 8.3B 的 mini 版;定位是端到端语音对话大模型(不是纯 TTS),流式依赖 vLLM 后端;GPU 需求官方未公布,是选型中的不确定项。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub(Step-Audio 2) | https://github.com/stepfun-ai/Step-Audio2 | [shields-verified] |
| Star 数 | 1.5k | [shields-verified] |
| License | Apache-2.0 | [shields-verified] + [正文已抓取] https://cdn.jsdelivr.net/gh/stepfun-ai/Step-Audio2@main/LICENSE(11357 B,Apache 2.0 全文) |
| Last commit | march | [shields-verified] |
| ⚠️ 旧仓库 Step-Audio | https://github.com/stepfun-ai/Step-Audio(42 stars,Apache-2.0,last commit march)—— README 首行明确写着:"This repository is no longer maintained, please refer to: Step-Audio2 / Step-Audio-R1 / Step-Audio-EditX" | [正文已抓取] https://cdn.jsdelivr.net/gh/stepfun-ai/Step-Audio@main/README.md(37303 B) |
| 官方 GitHub(语音推理) | https://github.com/stepfun-ai/Step-Audio-R1 — 978 stars?(见下) | 见 §4.4 |
| 官方 GitHub(音频编辑) | https://github.com/stepfun-ai/Step-Audio-EditX — 978 stars,Apache-2.0,last commit april | [shields-verified] |
| HuggingFace | https://huggingface.co/stepfun-ai/Step-Audio-2-mini(另有 -Base、-Think) | [正文已抓取] https://hf-mirror.com/stepfun-ai/Step-Audio-2-mini/raw/main/README.md(200,32766 B);license tag = apache-2.0 |
| arXiv(Step-Audio 2) | https://arxiv.org/abs/2507.16632 = Step-Audio 2 Technical Report,2025-07-22(v3) | [arxiv-api-verified] |
| arXiv(Step-Audio 初代) | https://arxiv.org/abs/2502.11946 = Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction,2025-02-17 | [arxiv-api-verified] |
| arXiv(Step-Audio-R1) | https://arxiv.org/abs/2511.15848 = Step-Audio-R1 Technical Report,2025-11-19 | [arxiv-api-verified] |
| arXiv(Step-Audio-EditX) | https://arxiv.org/abs/2511.03601 = Step-Audio-EditX Technical Report,2025-11-05 | [arxiv-api-verified] |
| 项 | 值 | 来源 |
|---|---|---|
| 开源部分 | 仅 Step-Audio 2 mini(+ Base / Think 变体);完整版 Step-Audio 2 未开源(README:"We release Step-Audio 2 mini ... under Apache 2.0") | 【厂商自报】 |
| Step-Audio-2-mini 实际参数 | 8,315,179,264 ≈ 8.32B(BF16) | [正文已抓取] https://hf-mirror.com/api/models/stepfun-ai/Step-Audio-2-mini?blobs=false |
| Step-Audio 初代 | 130B 多模态模型(仅开源 Step-Audio-Chat 变体)+ Step-Audio-TTS-3B + Step-Audio-Tokenizer | 【厂商自报】Step-Audio README |
| VRAM | 未查到(官方未公布;vLLM docker 参数未给显存要求) | — |
⚠️ Step-Audio 2 的评测表主要是 ASR(语音识别)指标,不是 TTS 合成指标。 README 原文:"CER for Chinese, Cantonese and Japanese and WER for Arabian and English"。
ASR CER/WER(【厂商自报】,README 评测表):
| 测试集 | Step-Audio 2 | Step-Audio 2 mini | Kimi-Audio | Qwen-Omni |
|---|---|---|---|---|
| AISHELL | 0.63 | 0.78 | 0.64 | 1.17 |
| AISHELL-2 | 2.10 | 2.16 | 2.67 | 2.40 |
| FLEURS Chinese | 2.68 | 2.53 | 2.91 | 7.01 |
| KeSpeech phase1 | 3.63 | 3.97 | 5.11 | 6.45 |
| WenetSpeech meeting | 4.75 | 4.87 | 5.21 | 6.61 |
| WenetSpeech net | 4.67 | 4.82 | 5.93 | 5.24 |
| 中文平均 | 3.08 | 3.19 | 3.75 | 4.81 |
TTS 侧 Seed-TTS test-zh CER / SIM:未查到(该模型不以 SeedTTS 零样本克隆为主指标,README 无此表)。 UTMOS / 自然度:未查到。
| 项 | 值 | 来源 |
|---|---|---|
| 流式支持 | 支持(需 vLLM 后端) — README:"We highly recommend using our vLLM backend for faster and streaming inference" | 【厂商自报】 |
| vLLM 镜像 | stepfun2025/vllm:step-audio-2-v20250909,--max-model-len 16384 --max-num-seqs 32 | 【厂商自报】 |
| TTFB / 首包延迟(ms) | 未查到(README 无 ms 数字) | — |
| RTF | 未查到 | — |
环境要求【厂商自报】:Python ≥ 3.10;PyTorch ≥ 2.3-cu121;transformers==4.49.0(版本锁定);pip install torchaudio librosa onnxruntime s3tokenizer diffusers hyperpyyaml。 vLLM 自定义分支:https://github.com/stepfun-ai/vllm/tree/step-audio2-mini(非上游)。
GitHub issue 检索:
| Issue | 标题 | 状态 |
|---|---|---|
| #57 | GPU Requirements for Matching StepFun Console Performance with vLLM? | open → 说明官方未给出清晰 GPU 需求 |
| #66 | vllm环境 | open |
| #61 | undefined symbol: ncclCommInitRankScalable | open |
| #86 | Step-Audio-2-mini running on AMD RDNA 3.5 (gfx1151 / Strix Halo) via ROCm — build fixes and Dockerfile | open |
| #68 | vLLM Prefill Performance: torch.compile Shows No Significant Improvement | open |
| #69 | How to Support Longer Audio in the vLLM Implementation? | open |
sm_120 / Blackwell 专项检索:未发现针对性报告(检索到的是 vLLM/ROCm 类通用问题)。
部分适合,但不是首选。 理由:
结论速览:11k stars、0.5B、中英双语零样本克隆。代码 Apache-2.0,但 HuggingFace 权重是 cc-by-nc-sa-4.0(禁止商用) → 整体判定:不可商用(除非只用代码+自训权重)。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/SparkAudio/Spark-TTS | [shields-verified] |
| Star 数 | 11k | [shields-verified] |
| License(仓库 shields) | Apache-2.0 | [shields-verified] |
| License(仓库 LICENSE 文件) | Apache-2.0(11357 B 全文) | [正文已抓取] https://cdn.jsdelivr.net/gh/SparkAudio/Spark-TTS@main/LICENSE |
| Last commit | april 2025 | [shields-verified] |
| HuggingFace | https://huggingface.co/SparkAudio/Spark-TTS-0.5B | [正文已抓取] https://hf-mirror.com/SparkAudio/Spark-TTS-0.5B/raw/main/README.md(200,6467 B) |
| ⚠️ HF license | cc-by-nc-sa-4.0 ← 禁止商用 | [正文已抓取] 模型卡 frontmatter license: cc-by-nc-sa-4.0;且 https://hf-mirror.com/api/models/SparkAudio/Spark-TTS-0.5B?blobs=false 的 tags 含 "license:cc-by-nc-sa-4.0" |
| arXiv | https://arxiv.org/abs/2503.01710 = Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens,2025-03-03 | [arxiv-api-verified] |
| 资产 | 许可 | 可否商用 |
|---|---|---|
| 推理代码(GitHub) | Apache-2.0 | ✅ 可 |
预训练权重(HuggingFace SparkAudio/Spark-TTS-0.5B) | CC-BY-NC-SA-4.0 | ❌ 不可(NonCommercial 明确禁止商业目的) |
判定:直接拿官方权重做商业数字人 → 不允许。 只有「用 Apache-2.0 的代码 + 完全自训权重」才绕开,但那等于放弃官方检查点。任何"看 GitHub 徽章写 Apache 就以为能商用"的做法在此模型上会踩雷。Spark-TTS 主 README 与 HF 模型卡互相矛盾,且 README 未对此做任何说明。
| 项 | 值 | 来源 |
|---|---|---|
| 参数量 | 0.5B(模型名即 Spark-TTS-0.5B) | 【厂商自报】模型名 + 第三方表 |
| 架构 | 完全基于 Qwen2.5,无独立 flow-matching 模型;BiCodec 单流解耦语音 token(语义 token + 固定长度全局说话人 token) | 【厂商自报】README / 论文摘要 |
| 训练数据 | VoxBox,100,000 小时带属性标注 | 【厂商自报】论文摘要 |
| VRAM | 未查到(README 未给 GB 数;0.5B 体量本身很小) | — |
【第三方表】VoxCPM2 README 与 MOSS-TTS README 两表一致:
👉 SIM 66.0 是本文所有模型中最低之一(对比 Qwen3-TTS 77.0、VoxCPM2 79.5、LongCat 81.8)——音色相似度是 Spark-TTS 的短板。 UTMOS / 自然度:未查到。
| 项 | 值 | 来源 |
|---|---|---|
| 流式支持 | 未声明支持流式合成(无 streaming API) | 【厂商自报】README 无 streaming 章节 |
| 生产部署 | 提供 NVIDIA Triton + TensorRT-LLM 参考实现 | 【厂商自报】README Runtime 章节 |
| 平均延迟 / RTF(L20 单卡,26 组 prompt/text,合计 169 秒音频) | 并发 1:Avg Latency 876.24 ms,RTF 0.1362;并发 2:920.97 ms,RTF 0.0737;并发 4:1611.51 ms,RTF 0.0704 | 【厂商自报】README 基准表(⚠️ 这是平均延迟,不是首包延迟 TTFB) |
| TTFB | 未查到(未单独给出) | — |
python=3.12;PyTorch 2.5+;pip install -r requirements.txt。不适合作为首选。 理由:
结论速览:仓库 boson-ai/higgs-audio 的 README 现在首屏就在推销 Higgs Audio v3,并明确 v3 采用 Research and Non-Commercial License。v2 文档被归档到 README_V2.md。v2 权重 HF license 是 other(语义不明),v2/v3 都没有 arXiv 论文。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/boson-ai/higgs-audio | [shields-verified] |
| Star 数 | 8.4k | [shields-verified] |
| License(仓库) | Apache-2.0 | [shields-verified] + [正文已抓取] https://cdn.jsdelivr.net/gh/boson-ai/higgs-audio@main/LICENSE(10141 B,Apache 2.0 全文) |
| Last commit | june | [shields-verified] |
| ⚠️ README 现状 | 主 README 已改为 Higgs Audio v3 的推广页,原文:"🎉 Higgs Audio v3 is here — you no longer need this repo!"、"Don't clone this repository to use the latest model.";v3 许可 = Boson Higgs Audio v3 Research and Non-Commercial License("Production / hosted / revenue-generating use requires a separate commercial license");并明确提示 v3 不依赖本仓库代码 | [正文已抓取] https://cdn.jsdelivr.net/gh/boson-ai/higgs-audio@main/README.md(3357 B) |
| v2 文档位置 | README_V2.md(归档说明页) | [正文已抓取] https://cdn.jsdelivr.net/gh/boson-ai/higgs-audio@main/README_V2.md(17233 B) |
| HuggingFace(v2 生成模型) | https://huggingface.co/bosonai/higgs-audio-v2-generation-3B-base | [正文已抓取] https://hf-mirror.com/bosonai/higgs-audio-v2-generation-3B-base/raw/main/README.md(200,25825 B) |
| ⚠️ HF license(v2) | license: other(未指明具体协议) | [正文已抓取] 模型卡 frontmatter + https://hf-mirror.com/api/models/bosonai/higgs-audio-v2-generation-3B-base?blobs=false tags = license:other |
| 姊妹权重 | bosonai/higgs-audio-v2-tokenizer(README 要求与该 tokenizer 配套加载) | 【厂商自报】README_V2 |
| v3 权重 | https://huggingface.co/bosonai/higgs-audio-v3-tts-4b(非商用许可) | 【厂商自报】主 README |
| arXiv 论文 | 未查到 ❌ —— 对 all:"Higgs Audio"、ti:"Higgs"、all:"Boson AI" 多次检索 arXiv API,均无 Higgs Audio 的论文条目。官方只提供博客:https://boson.ai/blog/higgs-audio-v2、https://www.boson.ai/blog/higgs-audio-v2.5 | [arxiv-api-verified](检索为空) |
| 项 | 值 | 来源 |
|---|---|---|
| 参数量(v2) | 3B(README 徽章原文:"3.6B LLM + 2.2B audio adapter") | 【厂商自报】README_V2 |
| HF 实际张量总量 | 5,771,283,456 ≈ 5.77B(BF16) | [正文已抓取] https://hf-mirror.com/api/models/bosonai/higgs-audio-v2-generation-3B-base?blobs=false |
| v2.5 | 压缩到 1B 参数,速度/精度超过 3B 版(GRPO 对齐 + Voice Bank 数据) | 【厂商自报】README_V2 |
| 训练数据 | 超过 1000 万小时音频 | 【厂商自报】README_V2 |
| VRAM | README 原文:"For optimal performance, run the generation examples on a machine equipped with GPU with at least 24GB memory!" | 【厂商自报】README_V2 |
Seed-TTS-eval(【第三方表】,VoxCPM2 README 与 MOSS-TTS README 两表一致):
自然度/表现力(【厂商自报】):在 EmergentTTS-Eval 上对 gpt-4o-mini-tts 的胜率 75.7%(Emotions 类) 与 55.7%(Questions 类)。 UTMOS:未查到。
| 项 | 值 | 来源 |
|---|---|---|
| 流式支持 | 支持(HiggsAudioServeEngine + vLLM 后端,提供 OpenAI 兼容 API server) | 【厂商自报】README_V2(examples/vllm) |
| TTFB | ⚠️ ~650 ms("Getting ttfb ~650ms using vllm") | 【用户报告】 GitHub issue #37(https://api.github.com/search/issues?q=repo:boson-ai/higgs-audio+...)—— 非官方数字 |
| RTF | 未查到(官方无 RTF) | — |
| 其他实测 | 【用户报告】issue #186:社区 C++/GGML 实现 "8.8–10.1× real-time (warmed) and 8.5× real-time (long)" | 用户自建实现,非官方 |
环境要求【厂商自报】:推荐 NVIDIA 深度学习容器 nvcr.io/nvidia/pytorch:25.02-py3 或 25.01-py3;pip install -r requirements.txt && pip install -e .;支持 vllm 后端。
GitHub issue 检索(repo:boson-ai/higgs-audio):
| Issue | 标题 | 状态 | 说明 |
|---|---|---|---|
| #39 | Blackwell Support on Docker | open | ⚠️ sm_120/Blackwell 官方支持仍缺 |
| #177 | OSError: libcudart.so.13: cannot open shared object file | open | ⚠️ CUDA 13 环境下直接加载失败 |
| #37 | Getting ttfb ~650ms using vllm | open | 首包延迟实测 ~650ms |
| #18 | vllm does not work | open | vLLM 路径不稳定 |
| #186 | [audio.cpp] Portable C++/GGML implementation … 8.8–10.1× real-time | open | 社区替代实现 |
不建议作为商业数字人底座。 理由:
license: other(含义不明,不能推定可商用);同一仓库主线已切到明确非商用的 v3;且 v3 不再使用本仓库代码,意味着 v2 路径已进入维护末期。libcudart.so.13 加载失败 issue #177 → 本机 CUDA 13.2 / RTX 5060 Ti 环境风险很高。结论速览:54k stars 的明星仓库,MIT 代码许可,但微软已于 2025-09-05 移除 VibeVoice-TTS 代码("widespread misuse"),官方模型表把 VibeVoice-TTS-1.5B 标为 Disabled。当前可用的是 VibeVoice-Realtime-0.5B(流式 TTS,~200ms 首音,仅英文单说话人)。官方 README 明确写"不建议用于商业或真实场景"。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/microsoft/VibeVoice | [shields-verified] |
| Star 数 | 54k | [shields-verified] |
| License | MIT | [shields-verified] + [正文已抓取] https://cdn.jsdelivr.net/gh/microsoft/VibeVoice@main/LICENSE(1066 B,MIT 全文,Copyright (c) 2025 Microsoft) |
| Last commit | september | [shields-verified] |
| HuggingFace | https://huggingface.co/microsoft/VibeVoice-1.5B(TTS,Disabled);https://huggingface.co/microsoft/VibeVoice-Realtime-0.5B(流式 TTS,可用);https://huggingface.co/microsoft/VibeVoice-ASR(ASR 7B) | [正文已抓取] https://hf-mirror.com/microsoft/VibeVoice-1.5B/raw/main/README.md(200,7279 B)、.../VibeVoice-Realtime-0.5B/raw/main/README.md(200,10160 B) |
| HF license tag | mit(1.5B 与 Realtime-0.5B 均为 mit) | [正文已抓取] https://hf-mirror.com/api/models/microsoft/VibeVoice-1.5B?blobs=false |
| arXiv(VibeVoice) | https://arxiv.org/abs/2508.19205 = VibeVoice Technical Report,2025-08-26 | [arxiv-api-verified] |
| 其他相关论文 | VibeVoice-ASR 报告 https://arxiv.org/abs/2601.18184;ASR-Streaming https://arxiv.org/abs/2609.02812;BitNet https://arxiv.org/abs/2607.21075(均见 README 徽章,未逐一 API 核验) | 【厂商自报】 |
| TTS 报告另发 | README 的 "TTS Report" 徽章指向 OpenReview https://openreview.net/pdf?id=FihSkzyxdv(ICLR 2026 Oral),并非 arXiv | 【厂商自报】README |
| 模型 | 上下文 | 生成长度 | 实际张量数 | 来源 |
|---|---|---|---|---|
| VibeVoice-1.5B | 64K | ~90 分钟 | 2,704,021,985 ≈ 2.70B(BF16) | [正文已抓取] HF API |
| VibeVoice-Large | 32K | ~45 分钟 | 权重 Disabled(未发布) | 【厂商自报】vv_tts.md |
| VibeVoice-Realtime-0.5B | 8k(~10 分钟音频) | 单说话人 | 1,017,626,722 ≈ 1.02B(BF16) | [正文已抓取] HF API |
| VibeVoice-ASR-7B | 64K | 60 分钟单次 | — | 【厂商自报】 |
| 架构 | Qwen2.5 骨干 + 7.5 Hz 连续语音 tokenizer(Acoustic + Semantic)+ 扩散头 | — | — | 【厂商自报】vv_tts.md |
| Tokenizer 细节 | Acoustic tokenizer 对 24kHz 输入 3200× 下采样(Encodec 压缩率提升 80×);encoder/decoder 各 ~340M;扩散头 ~123M(4 层) | — | — | 【厂商自报】HF VibeVoice-1.5B 模型卡 |
Seed-TTS-eval(【第三方表】):
VibeVoice-Realtime-0.5B 官方评测(【厂商自报】docs/vibevoice-realtime-0.5b.md):
| 测试集 | 模型 | WER ↓ | Speaker Sim ↑ |
|---|---|---|---|
| LibriSpeech test-clean | VibeVoice-Realtime-0.5B | 2.00 | 0.695 |
| LibriSpeech test-clean | VALL-E 2 | 2.40 | 0.643 |
| LibriSpeech test-clean | Voicebox | 1.90 | 0.662 |
| SEED test-en | VibeVoice-Realtime-0.5B | 2.05 | 0.633 |
| SEED test-en | Seed-TTS | 2.25 | 0.762 |
⚠️ 注意:官方只给英文(LibriSpeech / SEED test-en)数字,无中文指标 —— 因为 Realtime 版仅英文。
UTMOS / 自然度:未查到。
| 项 | 值 | 来源 |
|---|---|---|
| VibeVoice-TTS(1.5B) | 非流式,90 分钟单次整段生成 | 【厂商自报】 |
| VibeVoice-Realtime-0.5B | 流式(streaming text input + 增量扩散) | 【厂商自报】 |
| 首音延迟(首次可听) | ⚠️ 两个官方数字不一致:<br>• docs/vibevoice-realtime-0.5b.md:~200 milliseconds("It produces initial audible speech in ~200 milliseconds (hardware dependent)")<br>• 主 README.md 第 184 行:~300 milliseconds("Real-time TTS (~300 milliseconds first audible latency)")<br>• 同文档另注:"Due to network latency, the time when audio playback is heard may exceed the ~300 ms first speech chunk generation latency" | 【厂商自报】两处冲突,取保守值 300ms |
| 硬件基线 | "NVIDIA T4 / Mac M4 Pro achieve real-time performance in our tests" | 【厂商自报】 |
| 上下文 | 8k context(~10 分钟音频) | 【厂商自报】 |
pip install flash-attn --no-build-isolation);pip install -e .[streamingtts]。docs/vibevoice-tts.md 的 Installation 章节直接写 "Disabled due to widespread misuse."GitHub issue 检索:
| Issue | 标题 | 状态 |
|---|---|---|
| #210 | CUDA OOM with just 1 minutes audio | open |
| #367 | [Docs] Document the VRAM-vs-audio-duration relationship — RTX 4090 (24GB) OOMs on >30min audio at default sdpa attention | open |
| #377 | docs(asr): document VRAM-vs-audio-duration limits and add chunked inference | open |
| #332 | Token indices sequence length is longer than the specified maximum (153482 > 131072) | open |
| #185 | [Performance] Slow Inference on AMD 7900 XTX (Windows/ROCm) — ~1.0s/it | open |
| #340 | NVIDIA DGX Spark Support | open |
| #418 | CPU-only mode with SDPA attention by default | closed |
| #1? | repo:microsoft/VibeVoice+5090 检索 → total_count = 1(#210 CUDA OOM) | — |
不推荐。 理由:
VibeVoice-TTS-1.5B 已 Disabled。结论速览:15k stars、Apache-2.0、1B(+ Llama-3.2-1B 骨干)、对话式语音生成。官方自述非英语能力很弱("likely won't do well"),无中文指标,无 arXiv 论文。不适合中文数字人。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/SesameAILabs/csm | [shields-verified](且 HTML 抓取偶发 200) |
| Star 数 | 15k | [shields-verified] |
| License | Apache-2.0 | [shields-verified] + [正文已抓取] https://cdn.jsdelivr.net/gh/SesameAILabs/csm@main/LICENSE(11357 B,Apache 2.0 全文) |
| Last commit | may 2025 | [shields-verified] |
| HuggingFace | https://huggingface.co/sesame/csm-1b | [正文已抓取](HTML 页面 200,49036 B)https://hf-mirror.com/sesame/csm-1b;模型卡 raw README 返回 401(仓库 gated: auto,需登录接受条款才能读 raw) |
| HF license tag | apache-2.0 | [正文已抓取] https://hf-mirror.com/api/models/sesame/csm-1b → tags 含 "license:apache-2.0","gated":"auto" |
| HF 官方文档 | https://huggingface.co/sesame/csm_1b(README 中列出的另一个 checkpoint 路径) | 【厂商自报】README |
| arXiv 论文 | 未查到 ❌ —— 对 all:"Sesame CSM conversational speech" 等检索 arXiv API 无条目。官方仅有博客:https://www.sesame.com/research/crossing_the_uncanny_valley_of_voice | [arxiv-api-verified](检索为空) |
| 架构 | Llama 骨干 + 小型音频解码器,产出 Mimi(kyutai/mimi)RVQ 音频码 | 【厂商自报】README |
| 项 | 值 | 来源 |
|---|---|---|
| 参数量 | 1B(CSM 变体);HF 实际张量总量 1,552,791,552 ≈ 1.55B(F32,即未 bf16 存储);另需 meta-llama/Llama-3.2-1B 作为骨干 | 【厂商自报】README + [正文已抓取] https://hf-mirror.com/api/models/sesame/csm-1b?blobs=false |
| 依赖门禁 | 需访问 meta-llama/Llama-3.2-1B 和 sesame/csm-1b 两个 gated 仓库(需 huggingface-cli login) | 【厂商自报】README Requirements |
| VRAM | 未查到 | — |
未查到 —— 且这是该模型的设计限制,不是数据缺失。 README FAQ 原文:
"Does it support other languages?" — "The model has some capacity for non-English languages due to data contamination in the training data, but it likely won't do well."
| 项 | 值 | 来源 |
|---|---|---|
| 流式支持 | 未声明(无 streaming API;generate 一次性返回整段 audio tensor,max_audio_length_ms=10_000) | 【厂商自报】README |
| TTFB / RTF | 未查到 | — |
环境要求【厂商自报】:CUDA GPU;已在 CUDA 12.4 和 12.6 上测试;Python 3.10 推荐;部分音频操作需要 ffmpeg;export NO_TORCH_COMPILE=1(禁用 Mimi 懒编译)。
GitHub issue 检索(repo:SesameAILabs/csm+blackwell+OR+5090+OR+sm_120):
| Issue | 标题 | 状态 |
|---|---|---|
| #118 | NVIDIA GeForce RTX 5080 with CUDA capability sm_120 is not compatible with the current PyTorch installation | closed |
triton 无法安装,需改用 pip install triton-windows。不适合中文数字人。 理由:
结论速览:Apache-2.0(GitHub LICENSE)但 HF 模型卡 frontmatter 写 license: mit —— 两者冲突,商用前必须书面确认。中文 CER 0.89(RL 版)/ 1.03(base),SIM 76.1~76.4,支持流式,1.5B,是 CosyVoice3 评测表中的对照模型。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/zai-org/GLM-TTS | [shields-verified] |
| Star 数 | 1.1k | [shields-verified] |
| License(shields) | Apache-2.0 | [shields-verified] |
| License(仓库 LICENSE 文件) | Apache-2.0(11338 B 全文) | [正文已抓取] https://cdn.jsdelivr.net/gh/zai-org/GLM-TTS@main/LICENSE |
| ⚠️ License(HF 模型卡) | license: mit ← 与仓库 LICENSE 冲突 | [正文已抓取] https://hf-mirror.com/zai-org/GLM-TTS/raw/main/README.md frontmatter 第 8 行 license: mit;HF API tags = license:mit |
| Last commit | april | [shields-verified] |
| HuggingFace | https://huggingface.co/zai-org/GLM-TTS | [正文已抓取] https://hf-mirror.com/zai-org/GLM-TTS/raw/main/README.md(200,4617 B) |
| 中文说明 | README_zh.md | [正文已抓取] https://cdn.jsdelivr.net/gh/zai-org/GLM-TTS@main/README_zh.md(13731 B) |
| ModelScope | https://modelscope.cn/models/ZhipuAI/GLM-TTS | 【厂商自报】README |
| 在线体验 | https://audio.z.ai/ | 【厂商自报】README |
| arXiv | https://arxiv.org/abs/2512.14291 = GLM-TTS Technical Report,2025-12-16 | [arxiv-api-verified] |
| 项 | 值 | 来源 |
|---|---|---|
| 参数量 | 1.5B | 【第三方表】VoxCPM2 README 与 MOSS-TTS README 两表一致记 GLM-TTS = 1.5B;MOSS 表另记 GLM-TTS-RL = 1.5B |
| HF 张量数 | safetensors: null(该仓库未提供 safetensors 元数据,无法从 API 取精确参数量) | [正文已抓取] https://hf-mirror.com/api/models/zai-org/GLM-TTS?blobs=false |
| 训练数据 | 仅 100k 小时(论文摘要,相对同代模型偏小) | 【厂商自报】论文摘要 |
| 架构 | 两阶段:Llama 架构 LLM 生成语音 token → Flow Matching 生成 mel → vocoder | 【厂商自报】README |
| 特色 | 音素级混合输入(Phoneme-in)解决多音字;GRPO 多奖励 RL;LoRA 声音定制 | 【厂商自报】README |
| VRAM | 未查到(README 未给 GB 数) | — |
Seed-TTS-eval(seed-tts-eval zh testset,未开 --phoneme):
| 模型 | CER ↓ | SIM ↑ | 开源 | 来源 |
|---|---|---|---|---|
| MegaTTS3 | 1.52 | 79.0 | 🔒 | 【厂商自报】GLM-TTS README |
| DiTAR | 1.02 | 75.3 | 🔒 | 同上 |
| CosyVoice3 | 1.12 | 78.1 | 🔒 | 同上 |
| Seed-TTS | 1.12 | 79.6 | 🔒 | 同上 |
| MiniMax | 0.83 | 78.3 | 🔒 | 同上 |
| CosyVoice2 | 1.38 | 75.7 | 👐 | 同上 |
| F5-TTS | 1.53 | 76.0 | 👐 | 同上 |
| FireRedTTS-2 | 1.14 | 73.6 | 👐 | 同上 |
| IndexTTS2 | 1.03 | 76.5 | 👐 | 同上 |
| VibeVoice | 1.16 | 74.4 | 👐 | 同上 |
| HiggsAudio-v2 | 1.50 | 74.0 | 👐 | 同上 |
| VoxCPM | 0.93 | 77.2 | 👐 | 同上 |
| GLM-TTS(Base) | 1.03 | 76.1 | 👐 | 【厂商自报】 |
| GLM-TTS_RL(Ours) | 0.89 | 76.4 | 👐 | 【厂商自报】 |
第三方交叉验证(【第三方表】MOSS-TTS README):GLM-TTS 1.5B → EN WER 2.23 / EN SIM 67.2 / ZH CER 1.03 / ZH SIM 76.1;GLM-TTS-RL → EN WER 1.91 / EN SIM 68.1 / ZH CER 0.89 / ZH SIM 76.4。 👉 自报与第三方完全吻合(1.03/76.1 与 0.89/76.4),可信度较高。
UTMOS / 自然度:未查到(README 未给 UTMOS)。
| 项 | 值 | 来源 |
|---|---|---|
| 流式支持 | 支持 — README Features:"Streaming Inference: Support real-time streaming audio generation, suitable for interactive applications";flow/flow.py 实现流式推理 | 【厂商自报】README |
| 推理命令 | python glmtts_inference.py --data=example_zh --exp_name=_test --use_cache | 【厂商自报】 |
| TTFB / RTF(ms) | 未查到 | — |
| 部署 | 提供 NPU(昇腾 CANN)与 GPU 两条路径;requirements_npu.txt | 【厂商自报】 |
环境要求【厂商自报】:Python 3.10–3.12;pip install -r requirements.txt;RL 模块需额外 clone s3prl 与 LaughterSegmentation 并下载 wavlm_large_finetune.pth。
GitHub issue 检索(repo:zai-org/GLM-TTS+error 与 +blackwell+OR+5090+OR+sm_120+OR+flash):
| Issue | 标题 | 状态 | 说明 |
|---|---|---|---|
| #9 | Is this project using float32? The GPU utilization on the 5060ti 16G is only 40% | open | ⚠️ RTX 5060 Ti 16GB(sm_120)上 GPU 利用率仅 40% —— 与本次选型目标硬件(RTX 5060 Ti 16GB)完全对应,是最直接的实测线索 |
| #51 | LLM inference speed: 27ms/token on A100 - how to optimize? | open | 速度优化空间 |
| #3 | requirements.txt conflicts | open | 依赖冲突 |
| #37 | No module named 'tn.chinese' | open | 中文前端依赖缺失 |
| #33 | TypeError: expected str, bytes or os.PathLike object, not NoneType | open | — |
| #1 | Update numpy version in requirements.txt | open | — |
| #44 | onnxruntime加速 | open | — |
| #54 | 基于昇腾NPU部署本模型时报错及运行时报错 | open | NPU 路径不稳 |
| #42 | win 和 macos 都安装 requirements 失败 | closed | — |
| #24 | Installation and Runtime Optimizations | closed | — |
有条件适合 —— 但许可证需先澄清。 理由:
requirements.txt 有已知冲突(#3)。注意:截至 2026-09-22,index-tts 仓库的最新版本是 IndexTTS-2.5,未发现 "IndexTTS-3"。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/index-tts/index-tts | [shields-verified] |
| Star 数 | 24k | [shields-verified] |
| License | bilibili Model Use License Agreement(自定义,shields 显示 "not identifiable by github") | [shields-verified] + [正文已抓取] https://cdn.jsdelivr.net/gh/index-tts/index-tts@main/LICENSE(10554 B,标题即 "bilibili Model Use License Agreement") |
| Last commit | august | [shields-verified] |
| HuggingFace | https://huggingface.co/IndexTeam/IndexTTS-2.5(另有 IndexTTS-2 / -1.5) | [正文已抓取] https://hf-mirror.com/IndexTeam/IndexTTS-2.5/raw/main/README.md(200,4272 B);license tag = license:other / bilibili-model-license |
| arXiv(2.5) | https://arxiv.org/abs/2601.03888 = IndexTTS 2.5 Technical Report,2026-01-07(v5) | [arxiv-api-verified] |
| arXiv(2) | https://arxiv.org/abs/2506.21619 = IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech,2025-06-23 | [arxiv-api-verified] |
| arXiv(1) | https://arxiv.org/abs/2502.05512 = IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System,2025-02-08 | [arxiv-api-verified] |
| 参数 | ~0.8B GPT 骨干(HF 卡明示);checkpoint:gpt.pth 3.26 GB、s2mel.pth 415 MB、codec.pth 607 MB(共 4.28 GB) | 【厂商自报】+ [正文已抓取] HF tree API |
| 语言 | 中、英、日、西、阿 | 【厂商自报】 |
| 输出 | 22.05 kHz | 【厂商自报】 |
| 中文质量 | ⚠️ 该数字是 IndexTTS2(1.5B)的,不是 2.5 的:test-zh CER 1.03 / SIM 76.5 | 【第三方表】VoxCPM2 README 行 IndexTTS2 | 1.5B、MOSS-TTS README 行 IndexTTS2 | 1.5B、GLM-TTS README 行 IndexTTS2 —— 三家表均标注为 IndexTTS2。IndexTTS-2.5 自己的 README/HF 卡未给 SeedTTS CER/SIM 数字 → IndexTTS-2.5 的中文指标:未查到 |
| VRAM | ~6 GB(HF 卡:"roughly 6 GB of VRAM for inference") | 【厂商自报】 |
| 流式 | 非原生流式 — issue #408 "How to stream Index-tts2 - Code / demo included"、issue #717 "感谢开源!能否有一些推理优化做到实时?我做的实时数字人急需!"(用户明确要实时数字人而不得) | 【用户报告】 |
| 商用 | ⚠️ 基本可商用,但有门槛:许可 2.2 条规定,若你或关联方月活 >1 亿 或 上一年营收 > 10 亿人民币,必须另行申请授权。另有 2c 条:不得用本模型改进其他 AI 模型(改进非商业 AI 模型除外)。 | [正文已抓取] LICENSE 第 2.2、2(c) 条 |
| 坑 | 需 CUDA Toolkit ≥12.8;issue #584 "Too long gpt_gen_time on RTX 5090 (Windows & WSL2)" ⚠️ sm_120 性能问题;issue #722 社区 C++/GGML 版在 RTX 5090 上达 3× 实时、长文本比 Python 快 5.6×;#728 Python 3.13;#304 deepspeed 安装失败 | 【用户报告】 |
实时数字人适配性:⚠️ 一般 —— 质量好(SIM 76.5)、显存低(6GB),但无原生流式,且社区反复提"要实时"需求;RTX 5090 上生成耗时过长(#584)。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/OpenMOSS/MOSS-TTS | [shields-verified] |
| Star 数 | 4.1k | [shields-verified] |
| License | Apache-2.0 | [shields-verified] + [正文已抓取] https://cdn.jsdelivr.net/gh/OpenMOSS/MOSS-TTS@main/LICENSE(11375 B,Apache 2.0 全文) |
| Last commit | september | [shields-verified] |
| HuggingFace | https://huggingface.co/OpenMOSS-Team/MOSS-TTS(8B)、.../MOSS-TTS-v1.5、.../MOSS-TTS-Local-Transformer-v1.5(4B)、.../MOSS-TTS-Realtime、.../MOSS-TTS-Nano | [正文已抓取] https://hf-mirror.com/OpenMOSS-Team/MOSS-TTS/raw/main/README.md(200,34624 B)、.../MOSS-TTS-Realtime/raw/main/README.md(200,9895 B);license tag = apache-2.0 |
| arXiv | https://arxiv.org/abs/2603.18090 = MOSS-TTS Technical Report,2026-03-18(v2);tokenizer 论文 https://arxiv.org/abs/2602.10934 = MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models,2026-02-11 | [arxiv-api-verified] |
| 参数 | MOSS-TTS (MossTTSDelay) = 8B(HF 实际 8,489,841,664 ≈ 8.49B BF16);MossTTSLocal = 1.7B(v1.5 升至 4B);MOSS-TTS-Nano = 0.1B;MOSS-TTS-Realtime = 未查到具体参数量 | [正文已抓取] HF API + 【厂商自报】README |
| 中文质量(Seed-TTS-eval) | MossTTSDelay 8B:ZH CER 1.37 / ZH SIM 76.98;MossTTSLocal 1.7B:ZH CER 1.44 / ZH SIM 79.62 | 【厂商自报】README 对比表 |
| 流式(MOSS-TTS-Realtime) | TTFB = 180 ms(after warm up);RTF = 0.51(单张 L20,SDPA + torch.compile);配 Qwen3.5-9B(vLLM)时 T_LLM首句(197ms) + TTFB(180ms) = 377 ms | 【厂商自报】README |
| VRAM | 8B 模型可跑在 8GB GPU 上(llama.cpp 路径优化;trt-8gb.yaml 低显存模式,分阶段加载) | 【厂商自报】README News(2026.3.10) |
| 商用 | Apache-2.0 → 可商用 | [正文已抓取] LICENSE |
| 服务化 | SGLang-Omni 与 vLLM-Omni 均 Day-0 支持,OpenAI 兼容 /v1/audio/speech | 【厂商自报】 |
| 坑 | 未发现 sm_120 专项 issue(本次未针对该仓库做完整 issue 挖掘) | — |
实时数字人适配性:✅ 强 —— TTFB 180ms、RTF 0.51、8B 可塞 8GB 显存、Apache-2.0 可商用、vLLM-Omni/SGLang-Omni 一等支持。建议纳入候选池。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/meituan-longcat/LongCat-AudioDiT | [shields-verified](README 亦 [正文已抓取] jsdelivr 9985 B) |
| Star 数 | 580 | [shields-verified] |
| License | MIT | [shields-verified] + [正文已抓取] https://cdn.jsdelivr.net/gh/meituan-longcat/LongCat-AudioDiT@main/LICENSE(1064 B,MIT,Copyright (c) 2026 Meituan) |
| Last commit | april | [shields-verified] |
| HuggingFace | https://huggingface.co/meituan-longcat/LongCat-AudioDiT-3.5B(另有 -1B) | [正文已抓取] https://hf-mirror.com/meituan-longcat/LongCat-AudioDiT-3.5B/raw/main/README.md(200,9756 B);license tag = mit |
| arXiv | https://arxiv.org/abs/2603.29339 = LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space,2026-03-31 | [arxiv-api-verified] |
| 参数 | 3.5B(HF 实际 3,833,985,217 ≈ 3.83B,F32 存储);另有 1B | [正文已抓取] HF API |
| 中文质量 | Seed-ZH SIM 0.818(自报 SOTA,超越 Seed-TTS 的 0.809);Seed-Hard SIM 0.797(vs 0.776) | 【厂商自报】论文摘要 + HF 卡 |
| 第三方口径 | test-ZH CER 1.09 / SIM 81.8;test-EN WER 1.50 / SIM 78.6;test-Hard CER 6.04 / SIM 79.7 | 【第三方表】VoxCPM2 README(SIM 81.8 是全表最高) |
| 商用 | MIT → 可商用 | [正文已抓取] LICENSE |
| 流式 | 非自回归扩散模型 → README 未声明流式支持;TTFB/RTF 未查到 | — |
| VRAM | 未查到(3.83B F32 存储,权重本体约 15GB+,需注意) | — |
| 坑 | 未做 sm_120 专项 issue 检索 | — |
实时数字人适配性:⚠️ 质量极强但实时性不明 —— SIM 81.8 为全场最高、MIT 可商用,但非自回归扩散架构通常不适合流式;F32 权重体积大。建议作为"离线高质量配音"候选,而非实时直播候选。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/fishaudio/fish-speech | [shields-verified] |
| Star 数 | 33k | [shields-verified] |
| License(shields) | "not identifiable by github" | [shields-verified] |
| ⚠️ License(LICENSE 文件) | FISH AUDIO RESEARCH LICENSE — 原文:"This Agreement is intended to allow research and non-commercial uses of the Materials free of charge. Any Commercial use of the Materials requires a separate license from Fish Audio." | [正文已抓取] https://cdn.jsdelivr.net/gh/fishaudio/fish-speech@main/LICENSE(10360 B,Last Updated: March 7, 2026) |
| Last commit | last wednesday(活跃) | [shields-verified] |
| HuggingFace | https://huggingface.co/fishaudio/s2-pro | [正文已抓取] https://hf-mirror.com/fishaudio/s2-pro/raw/main/README.md(200,5117 B) |
| ⚠️ HF 门禁 | license: other / license_name: fish-audio-research-license;门禁必勾选项原文:"I agree to use this model for non-commercial use ONLY" | [正文已抓取] 模型卡 frontmatter |
| arXiv | https://arxiv.org/abs/2603.08823 = Fish Audio S2 Technical Report,2026-03-09(v2) | [arxiv-api-verified] |
| 参数 | 4B(Slow AR)+ 400M(Fast AR);HF 实际张量 4,561,852,416 ≈ 4.56B(BF16) | 【厂商自报】HF 卡 + [正文已抓取] HF API |
| 中文质量 | Seed-TTS Eval 中文 WER(CER) = 0.54 %(原文标注 "best overall");英文 WER 0.99%(best overall) | 【厂商自报】fish-speech README |
| 第三方中文口径 | test-ZH CER 0.54(VoxCPM2 表);CV3-eval zh 2.65(VoxCPM2 表);MiniMax-MLS 中文 0.730(VoxCPM2 表) | 【第三方表】VoxCPM2 README |
| 流式 / 实时 | RTF 0.195(单张 NVIDIA H200);Time-to-first-audio ~100 ms;吞吐 3000+ acoustic tokens/s 且 RTF <0.5 | 【厂商自报】HF 卡 "Production Streaming Performance" |
| 编码 | RVQ 10 codebooks,~21 Hz 帧率 | 【厂商自报】 |
| 语言 | 80+ 语言(Tier 1: ja/en/zh) | 【厂商自报】 |
| 商用结论 | ❌ 不可商用(需单独商业授权,联系 business@fish.audio) | [正文已抓取] LICENSE + HF 门禁 |
| 坑 | 未做 sm_120 专项 issue 检索 | — |
实时数字人适配性:技术上最强(中文 CER 0.54、TTFA 100ms、RTF 0.195),但法律上出局(禁商用 + 门禁强制勾选非商用)。除非购买商业授权,否则不能用于商业数字人。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/k2-fsa/ZipVoice | [shields-verified] |
| Star 数 | 1.1k | [shields-verified] |
| License | Apache-2.0 | [shields-verified] + [正文已抓取] https://cdn.jsdelivr.net/gh/k2-fsa/ZipVoice@master/LICENSE(11357 B,Apache 2.0 全文) |
| Last commit | december 2025 | [shields-verified] |
| HuggingFace | https://huggingface.co/k2-fsa/ZipVoice(含 zipvoice / zipvoice_distill / zipvoice_dialog / zipvoice_dialog_stereo / libritts 变体) | [正文已抓取] https://hf-mirror.com/k2-fsa/ZipVoice/raw/main/README.md(200,3216 B) |
| arXiv | https://arxiv.org/abs/2506.13053 = ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching,2025-06-16(v3);对话版 https://arxiv.org/abs/2507.09318 = ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching,2025-07-12 | [arxiv-api-verified] |
| 参数 | 仅 123M(README 明示 "small and fast: only 123M parameters") | 【厂商自报】README |
| 语言 | 中、英 | 【厂商自报】 |
| 中文质量 | 未查到 —— ZipVoice README 与 HF 卡均无 SeedTTS CER/SIM 表(论文里有,本次未解析 PDF) | — |
| 速度(论文摘要) | 与 SOTA 质量相当,比 DiT-based flow-matching 基线小 3 倍、快至 30 倍;蒸馏版可将采样步降到 4 | 【厂商自报】论文摘要 / README |
| 流式 | 未声明流式(非自回归 flow matching,批次推理);支持 ONNX / TensorRT(GPU 上约 2× 吞吐) | 【厂商自报】README |
| 商用 | Apache-2.0 → 可商用 | — |
| 坑 | k2 依赖与 PyTorch/CUDA 版本严格绑定(e.g. k2==1.24.4.dev20250208+cuda12.1.torch2.5.1);⚠️ sm_120/CUDA 13 环境大概率没有预编译 k2 —— 官方 k2 索引以 cu12.x 为主 | 【厂商自报】README |
| 其他 | 支持中文多音字手工标注 <chang2>;ONNX INT8 CPU 路径;不要用 ONNX 跑 GPU(比 PyTorch 慢) | 【厂商自报】README |
实时数字人适配性:⚠️ 中等 —— 123M 极小、Apache-2.0 可商用、中英双语、有 TensorRT 路径;但非自回归、无流式 API,且 k2 编译在 CUDA 13/sm_120 上风险高。适合做轻量离线/准实时配音。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/zai-org/GLM-4-Voice(同仓库 https://github.com/THUDM/GLM-4-Voice) | [shields-verified] |
| Star 数 | 3.2k | [shields-verified] |
| License | Apache-2.0 | [shields-verified] |
| Last commit | december 2024 ⚠️(最旧) | [shields-verified] |
| HuggingFace | zai-org/glm-4-voice-9b、zai-org/glm-4-voice-tokenizer、zai-org/glm-4-voice-decoder | [正文已抓取] https://hf-mirror.com/zai-org/glm-4-voice-9b/raw/main/README.md(200,1143 B);HF API 显示 glm-4-voice-9b 下载 9124、likes 121 |
| arXiv | https://arxiv.org/abs/2412.02612 = GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot,2024-12-03 | [arxiv-api-verified] |
| 参数量 | 9B(模型名 glm-4-voice-9b) | 【厂商自报】 |
| 中文质量(SeedTTS CER/SIM) | 未查到 | — |
| 定位 | 端到端语音对话 Chatbot(非纯 TTS) | 【厂商自报】 |
实时数字人适配性:⚠️ 偏旧(2024-12 最后提交),非 2026 主流选型。
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/OpenBMB/MiniCPM-o | [shields-verified] |
| Star 数 | 26k | [shields-verified] |
| License | Apache-2.0 | [shields-verified] |
| Last commit | september | [shields-verified] |
| HuggingFace | openbmb/MiniCPM-o-2_6 | [正文已抓取] https://hf-mirror.com/openbmb/MiniCPM-o-2_6/raw/main/README.md(200,50293 B) |
| 定位 | 端到端多模态(视觉+语音)实时对话模型,含 TTS 能力但不是专用 TTS | 【厂商自报】 |
| 已知坑(README 自述) | ⚠️ torchcodec 与 CUDA 版本兼容问题:torch>=2.11 默认捆绑 CUDA 13.1,CUDA 12.x 环境会报 RuntimeError: Could not load libtorchcodec;workaround 是钉 cu128 或降 triton/torchcodec | [正文已抓取] README 第 202~210 行 |
| 中文质量(SeedTTS) | 未查到(非 SeedTTS 榜模型) | — |
| 项 | 值 | 核实 |
|---|---|---|
| 官方 GitHub | https://github.com/MoonshotAI/Kimi-Audio | [shields-verified] |
| Star 数 | 4.7k | [shields-verified] |
| License | "not specified"(⚠️ 仓库未挂明确许可,商用性不明) | [shields-verified] .../license/MoonshotAI/Kimi-Audio.json → not specified |
| Last commit | june 2025 | [shields-verified] |
| HuggingFace | moonshotai/Kimi-Audio-7B-Instruct | [正文已抓取] https://hf-mirror.com/moonshotai/Kimi-Audio-7B-Instruct/raw/main/README.md(200,7175 B) |
| arXiv | https://arxiv.org/abs/2504.18425 = Kimi-Audio Technical Report,2025-04-25 | [arxiv-api-verified] |
| 中文 ASR 表现 | AISHELL 0.64、AISHELL-2 2.67(见 Step-Audio 2 README 对比表,为 ASR 指标)【第三方表】 | — |
| 定位 | 音频理解 + 生成(非纯 TTS) | — |
| 模型 | 开源状态 | 依据 |
|---|---|---|
| Qwen3-TTS-Flash | ❌ 闭源,仅 API(阿里云百炼 Model Studio 产品) | 无开源仓库与权重;Qwen 官方开源的是 Qwen3-TTS 0.6B/1.7B 系列(§2)。Qwen3-TTS README 同时提供 DashScope API 链接 https://help.aliyun.com/zh/model-studio/qwen-tts-realtime【厂商自报】 |
| MiniMax Speech | ❌ 未开源 | github.com/MiniMax-AI/MiniMax-Speech → shields 返回 "repo not found" [shields-verified];论文存在:https://arxiv.org/abs/2505.07916 = MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder,2025-05-12 [arxiv-api-verified] |
| Qwen3-Omni | ✅ 开源,Apache-2.0,4k stars,last commit april | [shields-verified] https://github.com/QwenLM/Qwen3-Omni;SeedTTS test-zh CER 1.07(VoxCPM2 表)【第三方表】 |
| FunAudioLLM/CosyVoice | ✅ 24k stars,Apache-2.0,last commit may | [shields-verified];CosyVoice3 未开源(GLM-TTS 表标 🔒 No) |
| OpenBMB/MOSS-TTS | 见 §10.2 | — |
| 模型 | 代码许可 | 权重许可 | 可否商用 | 依据 |
|---|---|---|---|---|
| Qwen3-TTS | Apache-2.0 | Apache-2.0 (HF tag) | ✅ 可 | LICENSE 全文 + HF tag |
| VoxCPM2 | Apache-2.0 | Apache-2.0 | ✅ 可(README 明示 "free for commercial use") | LICENSE + HF tag + README 原文 |
| FireRedTTS-2 | Apache-2.0 | Apache-2.0 | ✅ 可 | LICENSE + HF tag |
| MOSS-TTS | Apache-2.0 | Apache-2.0 | ✅ 可 | LICENSE + HF tag |
| LongCat-AudioDiT | MIT | MIT | ✅ 可 | LICENSE + HF tag |
| ZipVoice | Apache-2.0 | Apache-2.0 (未在 HF 验证 tag,但仓库 Apache) | ✅ 可 | LICENSE |
| GLM-TTS | Apache-2.0 | ⚠️ HF 标 mit(与仓库冲突) | ⚠️ 大概率可,但文档冲突需确认 | 两侧 LICENSE 分别抓取 |
| Step-Audio 2 mini | Apache-2.0 | Apache-2.0 | ✅ 可 | LICENSE + HF tag |
| Sesame CSM | Apache-2.0 | Apache-2.0 | ✅ 可(但官方有 Misuse 条款禁止仿冒等用途) | LICENSE + HF tag |
| IndexTTS-2.5 | bilibili 自定义 | bilibili 自定义 | ⚠️ 可,但月活>1亿或营收>10亿人民币需另行授权;另禁止用于改进其他 AI 模型 | LICENSE 2.2 / 2(c) 条 |
| VibeVoice | MIT | MIT | ⚠️ 法律上可(MIT),但官方明文"不建议商用/仅限研发" | LICENSE(MIT) + README 声明 |
| Higgs Audio v2 | Apache-2.0(仓库) | ⚠️ license: other(含义不明) | ❌/⚠️ 不可推定可商用;主线 v3 已明确非商用 | LICENSE(Apache) vs HF tag(other) vs 主 README(v3 非商用) |
| Spark-TTS | Apache-2.0 | ⚠️ CC-BY-NC-SA-4.0 | ❌ 权重不可商用 | LICENSE(Apache) vs HF frontmatter(cc-by-nc-sa-4.0) |
| Fish Audio S2 Pro | Fish Audio Research License | Fish Audio Research License | ❌ 不可(商用需单独授权;HF 门禁强制勾"non-commercial ONLY") | LICENSE + HF gated fields |
| 模型 | 参数 | test-zh CER/WER % ↓ | test-zh SIM % ↑ | 数据来源 |
|---|---|---|---|---|
| LongCat-AudioDiT-3.5B | 3.5B | 1.09 | 81.8 | 【第三方表】VoxCPM2 README |
| VoxCPM2 | 2B | 0.97 | 79.5 | 【厂商自报】README |
| Qwen3-TTS-12Hz-1.7B-Base | 1.7B | 0.77(自报)/ 1.22~1.33(第三方) | 0.799(自报 SIM)/ 77.0~76.72(第三方) | 自报 vs 【第三方表】 |
| GLM-TTS_RL | 1.5B | 0.89 | 76.4 | 【厂商自报】+【第三方表】一致 |
| GLM-TTS (Base) | 1.5B | 1.03 | 76.1 | 【厂商自报】+【第三方表】一致 |
| IndexTTS2(⚠️ 非 2.5) | 1.5B | 1.03 | 76.5 | 【第三方表】三家一致(均标 IndexTTS2) |
| IndexTTS-2.5 | ~0.8B GPT 骨干 | 未查到 | 未查到 | 官方 README / HF 卡未给 SeedTTS 数字 |
| MossTTSLocal | 1.7B | 1.44 | 79.62 | 【厂商自报】 |
| MossTTSDelay | 8B | 1.37 | 76.98 | 【厂商自报】 |
| Fish Audio S2 Pro | 4B | 0.54 | 未查到 | 【厂商自报】README |
| FireRedTTS-2 | 1.5B | 1.14 | 73.6 | 【第三方表】三家一致 |
| VibeVoice-1.5B | 1.5B | 1.16 | 74.4 | 【第三方表】三家一致 |
| HiggsAudio-v2 | 3B | 1.50 | 74.0 | 【第三方表】三家一致 |
| Spark-TTS-0.5B | 0.5B | 1.54 | 66.0 | 【第三方表】两家一致 |
| Sesame CSM | 1B | 未查到 | 未查到 | 官方自述非英语"likely won't do well" |
| Step-Audio 2 | 8.32B (mini) | 未查到(TTS 侧) | 未查到 | 仅 ASR 指标 |
| ZipVoice | 123M | 未查到 | 未查到 | 论文含表,本轮未解析 |
| OpenAudio-s1-mini(参照) | 0.5B | 1.18 | 68.5 | 【第三方表】VoxCPM2 README |
| 模型 | 流式 | TTFB / 首包(ms) | RTF | 来源类型 |
|---|---|---|---|---|
| Qwen3-TTS | ✅ | 97(端到端合成延迟,官方) | 未查到 | 【厂商自报】 |
| Fish Audio S2 Pro | ✅ | ~100(TTFA,H200) | 0.195(H200) | 【厂商自报】 |
| FireRedTTS-2 | ✅ | 140(L20) | 未查到 | 【厂商自报】 |
| MOSS-TTS-Realtime | ✅ | 180(L20,after warm up) | 0.51(L20) | 【厂商自报】 |
| VoxCPM2 | ✅ | 未查到 | ~0.30(4090)/ ~0.13(Nano-VLLM) | 【厂商自报】 |
| VibeVoice-Realtime-0.5B | ✅ | ~200 / ~300(官方两处冲突;取保守 300) | 未查到 | 【厂商自报】 |
| Higgs Audio v2 | ✅ | ~650(用户实测 vLLM) | 未查到 | 【用户报告】issue #37 |
| GLM-TTS | ✅ | 未查到 | 未查到 | 【厂商自报】 |
| Spark-TTS | ❌ | — | 0.1362(L20,并发1);平均延迟 876 ms | 【厂商自报】 |
| VibeVoice-TTS-1.5B | ❌(长文整段) | — | 未查到 | — |
| LongCat-AudioDiT | ❌(非自回归扩散) | 未查到 | 未查到 | — |
| IndexTTS-2.5 | ❌(社区有 streaming 尝试,issue #408) | 未查到 | 未查到 | — |
| ZipVoice | ❌ | 未查到 | 未查到 | — |
| Sesame CSM | ❌ | 未查到 | 未查到 | — |
| Step-Audio 2 | ✅(需 vLLM 后端) | 未查到 | 未查到 | 【厂商自报】 |
| 模型 | 参数量 | 官方 VRAM | 权重存储 dtype | 依据 |
|---|---|---|---|---|
| MineCPM-o / 其他 | — | — | — | — |
| MossTTSDelay (MOSS-TTS) | 8.49B | 8 GB 可跑(llama.cpp 优化 / trt-8gb.yaml 分阶段加载) | BF16 | 【厂商自报】+ HF API |
| Higgs Audio v2 | 5.77B(HF)/ "3.6B LLM + 2.2B adapter" | ≥24 GB(README 明示) | BF16 | 【厂商自报】+ HF API |
| Step-Audio-2-mini | 8.32B | 未查到 | BF16 | HF API |
| Fish Audio S2 Pro | 4.56B(4B Slow AR + 400M Fast AR) | 未查到 | BF16 | HF API + 卡片 |
| LongCat-AudioDiT-3.5B | 3.83B | 未查到 | F32 | HF API |
| VibeVoice-1.5B | 2.70B | 未查到(issue:24GB 卡 >30min 音频 OOM) | BF16 | HF API + issue #367 |
| VoxCPM2 | 2.29B / 标称 2B | ~8 GB | BF16 | HF API + 卡片 |
| Qwen3-TTS-12Hz-1.7B-Base | 1.93B / 标称 1.7B | 未查到 | BF16 | HF API |
| IndexTTS-2.5 | ~0.8B GPT 骨干 | ~6 GB | checkpoint 混合 | 【厂商自报】 |
| GLM-TTS | 1.5B(第三方表) | 未查到 | 未查到 | 【第三方表】 |
| FireRedTTS-2 | 1.5B(第三方表) | 9 GB(bf16)/ 14 GB(fp32) | bf16 | 【厂商自报】+【第三方表】 |
| Sesame CSM | 1.55B(HF)/ 标称 1B | 未查到 | F32 | HF API |
| VibeVoice-Realtime-0.5B | 1.02B | 未查到 | BF16 | HF API |
| Spark-TTS-0.5B | 0.5B | 未查到 | 未查到 | — |
| ZipVoice | 123M | 未查到 | 未查到 | 【厂商自报】 |
| 模型 | sm_120 相关发现 | 严重度 |
|---|---|---|
| Qwen3-TTS | 专项检索 total_count = 0(无任何 sm_120 issue) | 🟢 低 |
| FireRedTTS-2 | 未检索到 sm_120 issue;但锁 torch 2.7.1 + cu126 | 🟡 中(需对齐 CUDA 版本) |
| MOSS-TTS | 未发现专项 issue;官方维护 llama.cpp CUDA 路径(含 flash_attn 开关) | 🟡 中(待实测) |
| GLM-TTS | ⚠️ issue #9:RTX 5060 Ti 16G 上 GPU 利用率仅 40%,用户质疑是否 float32 —— 与本次目标硬件完全一致 | 🔴 高 |
| VoxCPM2 | ⚠️ #250 Blackwell/sm_120 Docker 支持(open)、#326 Blackwell CUDA graph 音质劣化、#282 Blackwell 比 4090 慢、#102 4090 segfault | 🔴 高 |
| Higgs Audio v2 | ⚠️ #39 Blackwell Docker 支持(open)、#177 libcudart.so.13 加载失败(open) | 🔴 高 |
| Sesame CSM | ⚠️ #118 RTX 5080 sm_120 与当前 PyTorch 不兼容(closed) | 🔴 高 |
| IndexTTS-2.5 | ⚠️ #584 RTX 5090 生成耗时过长(Windows & WSL2);#722 社区 C++/GGML 版在 5090 达 3× 实时 | 🟠 中高 |
| Spark-TTS | 未检索到 sm_120 issue;但权重禁商用 + 1.5 年未更新 | ⚪(已被许可排除) |
| ZipVoice | ⚠️ k2 预编译包主要面向 cu12.x,CUDA 13/sm_120 大概率需自行编译 | 🟠 中高 |
| VibeVoice | 未检索到 sm_120 专项;但 #210/#367 VRAM-时长耦合 OOM | 🟠 中高 |
| 档位 | 模型 | 核心理由 | 主要风险 |
|---|---|---|---|
| 第一梯队 | Qwen3-TTS(0.6B 优先试) | 97ms 官方延迟;原生流式;Apache-2.0 可商用;sm_120 零 issue;中文 CER 0.77~1.33 / SIM 76.7~79.9 | VRAM 未公布 |
| 第一梯队 | FireRedTTS-2 | 140ms 首包;原生流式对话;9GB 显存;4 说话人;Apache-2.0 可商用 | torch/cu126 锁定;单句 CER 1.14 一般 |
| 第一梯队 | MOSS-TTS-Realtime | 180ms TTFB;RTF 0.51;Apache-2.0 可商用;vLLM-Omni/SGLang 一等支持 | 细节数据需补测 |
| 第二梯队 | VoxCPM2 | RTF 0.13~0.30;48kHz;~8GB;中文 CER 0.97 / SIM 79.5;Apache-2.0 可商用 | sm_120 多个 open issue(#250/#282/#326) |
| 第二梯队 | GLM-TTS | 流式;中文 CER 0.89;音素级控制;自报与第三方吻合 | license 冲突(Apache vs MIT);#9 直指 5060 Ti 利用率 40% |
| 第三梯队 | IndexTTS-2.5 | 中文 SIM 76.5;~6GB 显存 | 无原生流式;#584 RTX 5090 慢;许可有营收门槛 |
| 需授权才可考虑 | Fish Audio S2 Pro | 中文 CER 0.54(最强);TTFA 100ms;RTF 0.195 | 禁商用,需购买商业授权 |
| 排除(许可) | Spark-TTS | — | 权重 CC-BY-NC-SA-4.0 |
| 排除(架构/定位) | Higgs Audio v2 / VibeVoice / Sesame CSM / LongCat-AudioDiT / ZipVoice / Step-Audio 2 | 见各自章节 | 24GB 显存 / 下架+仅英文 / 中文弱 / 非流式 / k2 编译 / 显存未知 |
cdn.jsdelivr.net/gh/... 或 raw.githubusercontent.com 实际抓取正文)| # | URL | 用途 | 结果 |
|---|---|---|---|
| 1 | https://cdn.jsdelivr.net/gh/OpenBMB/VoxCPM@main/README.md | VoxCPM2 正文 | 200 / 42395 B |
| 2 | https://cdn.jsdelivr.net/gh/OpenBMB/VoxCPM@main/LICENSE | VoxCPM2 许可 | 200 / 11298 B (Apache-2.0) |
| 3 | https://cdn.jsdelivr.net/gh/QwenLM/Qwen3-TTS@main/README.md(亦经 raw) | Qwen3-TTS 正文 | 200 / 60501 B |
| 4 | https://cdn.jsdelivr.net/gh/QwenLM/Qwen3-TTS@main/LICENSE | Qwen3-TTS 许可 | 200 / 11343 B (Apache-2.0) |
| 5 | https://cdn.jsdelivr.net/gh/FireRedTeam/FireRedTTS2@main/README.md | FireRedTTS-2 正文 | 200 / 14583 B |
| 6 | https://cdn.jsdelivr.net/gh/FireRedTeam/FireRedTTS2@main/LICENSE | FireRedTTS-2 许可 | 200 / 11357 B (Apache-2.0) |
| 7 | https://cdn.jsdelivr.net/gh/stepfun-ai/Step-Audio2@main/README.md(亦经 raw) | Step-Audio 2 正文 | 200 / 36591 B |
| 8 | https://cdn.jsdelivr.net/gh/stepfun-ai/Step-Audio2@main/LICENSE | Step-Audio 2 许可 | 200 / 11357 B (Apache-2.0) |
| 9 | https://cdn.jsdelivr.net/gh/stepfun-ai/Step-Audio@main/README.md | Step-Audio(旧)正文 | 200 / 37303 B |
| 10 | https://cdn.jsdelivr.net/gh/stepfun-ai/Step-Audio-R1@main/README.md | Step-Audio-R1 正文 | 200 / 15962 B |
| 11 | https://cdn.jsdelivr.net/gh/stepfun-ai/Step-Audio-EditX@main/README.md | Step-Audio-EditX 正文 | 200 / 44435 B |
| 12 | https://cdn.jsdelivr.net/gh/SparkAudio/Spark-TTS@main/README.md(亦经 raw) | Spark-TTS 正文 | 200 / 11232 B |
| 13 | https://cdn.jsdelivr.net/gh/SparkAudio/Spark-TTS@main/LICENSE | Spark-TTS 代码许可 | 200 / 11357 B (Apache-2.0) |
| 14 | https://cdn.jsdelivr.net/gh/boson-ai/higgs-audio@main/README.md | Higgs 主 README(v3 非商用) | 200 / 3357 B |
| 15 | https://cdn.jsdelivr.net/gh/boson-ai/higgs-audio@main/README_V2.md | Higgs v2 归档文档 | 200 / 17233 B |
| 16 | https://cdn.jsdelivr.net/gh/boson-ai/higgs-audio@main/LICENSE | Higgs 仓库许可 | 200 / 10141 B (Apache-2.0) |
| 17 | https://cdn.jsdelivr.net/gh/microsoft/VibeVoice@main/README.md | VibeVoice 主 README | 200 / 13031 B |
| 18 | https://cdn.jsdelivr.net/gh/microsoft/VibeVoice@main/docs/vibevoice-tts.md | VibeVoice-TTS 文档(Disabled) | 200 / 8126 B |
| 19 | https://cdn.jsdelivr.net/gh/microsoft/VibeVoice@main/docs/vibevoice-realtime-0.5b.md | VibeVoice-Realtime 文档 | 200 / 8095 B |
| 20 | https://cdn.jsdelivr.net/gh/microsoft/VibeVoice@main/LICENSE | VibeVoice 许可 | 200 / 1066 B (MIT) |
| 21 | https://cdn.jsdelivr.net/gh/SesameAILabs/csm@main/README.md | CSM 正文 | 200 / 5515 B |
| 22 | https://cdn.jsdelivr.net/gh/SesameAILabs/csm@main/LICENSE | CSM 许可 | 200 / 11357 B (Apache-2.0) |
| 23 | https://cdn.jsdelivr.net/gh/zai-org/GLM-TTS@main/README.md | GLM-TTS 正文 | 200 / 15190 B |
| 24 | https://cdn.jsdelivr.net/gh/zai-org/GLM-TTS@main/README_zh.md | GLM-TTS 中文 | 200 / 13731 B |
| 25 | https://cdn.jsdelivr.net/gh/zai-org/GLM-TTS@main/LICENSE | GLM-TTS 仓库许可 | 200 / 11338 B (Apache-2.0) |
| 26 | https://cdn.jsdelivr.net/gh/zai-org/GLM-4-Voice@main/README.md(亦 THUDM/) | GLM-4-Voice 正文 | 200 / 7739 B |
| 27 | https://cdn.jsdelivr.net/gh/k2-fsa/ZipVoice@master/README.md | ZipVoice 正文 | 200 / 15330 B |
| 28 | https://cdn.jsdelivr.net/gh/k2-fsa/ZipVoice@master/LICENSE | ZipVoice 许可 | 200 / 11357 B (Apache-2.0) |
| 29 | https://cdn.jsdelivr.net/gh/index-tts/index-tts@main/README.md | IndexTTS 正文 | 200 / 25219 B |
| 30 | https://cdn.jsdelivr.net/gh/index-tts/index-tts@main/LICENSE | IndexTTS 许可 | 200 / 10554 B (bilibili Model Use License) |
| 31 | https://cdn.jsdelivr.net/gh/OpenMOSS/MOSS-TTS@main/README.md | MOSS-TTS 正文 | 200 / 57241 B |
| 32 | https://cdn.jsdelivr.net/gh/OpenMOSS/MOSS-TTS@main/LICENSE | MOSS-TTS 许可 | 200 / 11375 B (Apache-2.0) |
| 33 | https://cdn.jsdelivr.net/gh/meituan-longcat/LongCat-AudioDiT@main/README.md | LongCat 正文 | 200 / 9985 B |
| 34 | https://cdn.jsdelivr.net/gh/meituan-longcat/LongCat-AudioDiT@main/LICENSE | LongCat 许可 | 200 / 1064 B (MIT) |
| 35 | https://cdn.jsdelivr.net/gh/fishaudio/fish-speech@main/README.md | Fish Audio S2 正文 | 200 / 12065 B |
| 36 | https://cdn.jsdelivr.net/gh/fishaudio/fish-speech@main/LICENSE | Fish Audio 许可 | 200 / 10360 B (Fish Audio Research License,禁商用) |
| 37 | https://cdn.jsdelivr.net/gh/microsoft/VibeASR.cpp@main/README.md | VibeASR.cpp 正文 | 200 / 7844 B |
| 38 | https://cdn.jsdelivr.net/gh/MoonshotAI/Kimi-Audio@main/README.md | Kimi-Audio 正文 | 200 / 21856 B |
| 39 | https://cdn.jsdelivr.net/gh/OpenBMB/MiniCPM-o@main/README.md | MiniCPM-o 正文 | 200 / 122707 B |
| 40 | https://raw.githubusercontent.com/OpenBMB/VoxCPM/main/README.md | VoxCPM2(raw 路径验证) | 200 |
| URL 模板 | 用途 |
|---|---|
https://img.shields.io/github/stars/<owner>/<repo>.json | star 数 |
https://img.shields.io/github/license/<owner>/<repo>.json | 仓库许可 |
https://img.shields.io/github/last-commit/<owner>/<repo>.json | 最后提交 |
| 已查询仓库 | OpenBMB/VoxCPM, QwenLM/Qwen3-TTS, FireRedTeam/FireRedTTS2, FireRedTeam/FireRedTTS, stepfun-ai/Step-Audio2, stepfun-ai/Step-Audio, stepfun-ai/Step-Audio-EditX, SparkAudio/Spark-TTS, boson-ai/higgs-audio, microsoft/VibeVoice, microsoft/VibeASR.cpp, SesameAILabs/csm, zai-org/GLM-TTS, zai-org/GLM-4-Voice, THUDM/GLM-4-Voice, k2-fsa/ZipVoice, index-tts/index-tts, MoonshotAI/Kimi-Audio, OpenBMB/MiniCPM-o, OpenBMB/MiniCPM-o-2_6, fishaudio/fish-speech, MiniMax-AI/MiniMax-Speech, OpenMOSS/MOSS-TTS, meituan-longcat/LongCat-AudioDiT, QwenLM/Qwen3-Omni, FunAudioLLM/CosyVoice, zai-org/GLM-TTS-Flash, microsoft/VibeVoice-Realtime |
| # | URL | 结果 |
|---|---|---|
| 41 | https://hf-mirror.com/openbmb/VoxCPM2/raw/main/README.md | 200 / 7939 B |
| 42 | https://hf-mirror.com/OpenBMB/VoxCPM-0.5B/raw/main/README.md | 200 / 12719 B |
| 43 | https://hf-mirror.com/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice/raw/main/README.md | 200 / 57846 B |
| 44 | https://hf-mirror.com/Qwen/Qwen3-TTS-12Hz-1.7B-Base/raw/main/README.md | 200 / 57817 B |
| 45 | https://hf-mirror.com/Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice/raw/main/README.md | 200 / 3263 B |
| 46 | https://hf-mirror.com/FireRedTeam/FireRedTTS2/raw/main/README.md | 200 / 8883 B |
| 47 | https://hf-mirror.com/stepfun-ai/Step-Audio-2-mini/raw/main/README.md | 200 / 32766 B |
| 48 | https://hf-mirror.com/SparkAudio/Spark-TTS-0.5B/raw/main/README.md | 200 / 6467 B |
| 49 | https://hf-mirror.com/bosonai/higgs-audio-v2-generation-3B-base/raw/main/README.md | 200 / 25825 B |
| 50 | https://hf-mirror.com/microsoft/VibeVoice-1.5B/raw/main/README.md | 200 / 7279 B |
| 51 | https://hf-mirror.com/microsoft/VibeVoice-Realtime-0.5B/raw/main/README.md | 200 / 10160 B |
| 52 | https://hf-mirror.com/microsoft/VibeVoice-ASR/raw/main/README.md | 200 / 4345 B |
| 53 | https://hf-mirror.com/sesame/csm-1b (HTML 页面;raw README 返回 401 gated) | 200 / 49036 B |
| 54 | https://hf-mirror.com/zai-org/GLM-TTS/raw/main/README.md | 200 / 4617 B |
| 55 | https://hf-mirror.com/zai-org/glm-4-voice-9b/raw/main/README.md | 200 / 1143 B |
| 56 | https://hf-mirror.com/k2-fsa/ZipVoice/raw/main/README.md | 200 / 3216 B |
| 57 | https://hf-mirror.com/IndexTeam/IndexTTS-2.5/raw/main/README.md | 200 / 4272 B |
| 58 | https://hf-mirror.com/IndexTeam/IndexTTS-2/raw/main/README.md | 200 / 2511 B |
| 59 | https://hf-mirror.com/OpenMOSS-Team/MOSS-TTS/raw/main/README.md | 200 / 34624 B |
| 60 | https://hf-mirror.com/OpenMOSS-Team/MOSS-TTS-v1.5/raw/main/README.md | 200 / 13744 B |
| 61 | https://hf-mirror.com/OpenMOSS-Team/MOSS-TTS-Realtime/raw/main/README.md | 200 / 9895 B |
| 62 | https://hf-mirror.com/meituan-longcat/LongCat-AudioDiT-3.5B/raw/main/README.md | 200 / 9756 B |
| 63 | https://hf-mirror.com/fishaudio/s2-pro/raw/main/README.md | 200 / 5117 B |
| 64 | https://hf-mirror.com/moonshotai/Kimi-Audio-7B-Instruct/raw/main/README.md | 200 / 7175 B |
| 65 | https://hf-mirror.com/openbmb/MiniCPM-o-2_6/raw/main/README.md | 200 / 50293 B |
| # | URL | 用途 |
|---|---|---|
| 66 | https://hf-mirror.com/api/models?search=<kw>&limit=25 | 模型检索(Qwen3-TTS / Sesame / csm / GLM-4-Voice / VibeVoice / GLM-TTS / MOSS-TTS / openaudio / Fish-Audio / LongCat-AudioDiT) |
| 67 | https://hf-mirror.com/api/models/openbmb/VoxCPM2?blobs=false | VoxCPM2:BF16 2,290,004,544;license:apache-2.0 |
| 68 | https://hf-mirror.com/api/models/Qwen/Qwen3-TTS-12Hz-1.7B-Base?blobs=false | Qwen3-TTS:BF16 1,928,677,440 |
| 69 | https://hf-mirror.com/api/models/stepfun-ai/Step-Audio-2-mini?blobs=false | Step-Audio-2-mini:BF16 8,315,179,264 |
| 70 | https://hf-mirror.com/api/models/FireRedTeam/FireRedTTS2?blobs=false | license:apache-2.0;safetensors null |
| 71 | https://hf-mirror.com/api/models/zai-org/GLM-TTS?blobs=false | ⚠️ license:mit(与仓库 Apache-2.0 冲突) |
| 72 | https://hf-mirror.com/api/models/IndexTeam/IndexTTS-2.5?blobs=false | license:other |
| 73 | https://hf-mirror.com/api/models/meituan-longcat/LongCat-AudioDiT-3.5B?blobs=false | F32 3,833,985,217;license:mit |
| 74 | https://hf-mirror.com/api/models/OpenMOSS-Team/MOSS-TTS?blobs=false | BF16 8,489,841,664;license:apache-2.0 |
| 75 | https://hf-mirror.com/api/models/microsoft/VibeVoice-1.5B?blobs=false | BF16 2,704,021,985;license:mit |
| 76 | https://hf-mirror.com/api/models/microsoft/VibeVoice-Realtime-0.5B?blobs=false | BF16 1,017,626,722;license:mit |
| 77 | https://hf-mirror.com/api/models/sesame/csm-1b?blobs=false | F32 1,552,791,552;license:apache-2.0;gated:auto |
| 78 | https://hf-mirror.com/api/models/bosonai/higgs-audio-v2-generation-3B-base?blobs=false | BF16 5,771,283,456;⚠️ license:other |
| 79 | https://hf-mirror.com/api/models/fishaudio/s2-pro?blobs=false | BF16 4,561,852,416;⚠️ license:other |
| 80 | https://hf-mirror.com/api/models/SparkAudio/Spark-TTS-0.5B?blobs=false | ⚠️ license:cc-by-nc-sa-4.0 |
| 81 | https://hf-mirror.com/api/models/zai-org/GLM-TTS/tree/main | 文件列表(0 GB → 仓库结构特殊) |
| 82 | https://hf-mirror.com/api/models/FireRedTeam/FireRedTTS2/tree/main | llm_pretrain.pt 8.27GB / llm_posttrain.pt 8.27GB / codec.pt 4.30GB |
| 83 | https://hf-mirror.com/api/models/IndexTeam/IndexTTS-2.5/tree/main | gpt.pth 3.26GB / s2mel.pth 415MB / codec.pth 607MB |
https://export.arxiv.org/api/query?id_list=<id> 与 search_query=)| # | arXiv ID | 验证得到的标题 | 发布 |
|---|---|---|---|
| 84 | 2601.15621 | Qwen3-TTS Technical Report | 2026-01-22 |
| 85 | 2606.06928 | VoxCPM2 Technical Report | 2026-06-05 |
| 86 | 2509.24650 | VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning | 2025-09-29 |
| 87 | 2509.02020 | FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot | 2025-09-02 |
| 88 | 2507.16632 | Step-Audio 2 Technical Report | 2025-07-22 |
| 89 | 2502.11946 | Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction | 2025-02-17 |
| 90 | 2511.15848 | Step-Audio-R1 Technical Report | 2025-11-19 |
| 91 | 2511.03601 | Step-Audio-EditX Technical Report | 2025-11-05 |
| 92 | 2503.01710 | Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens | 2025-03-03 |
| 93 | 2508.19205 | VibeVoice Technical Report | 2025-08-26 |
| 94 | 2512.14291 | GLM-TTS Technical Report | 2025-12-16 |
| 95 | 2412.02612 | GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot | 2024-12-03 |
| 96 | 2506.13053 | ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching | 2025-06-16 |
| 97 | 2507.09318 | ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching | 2025-07-12 |
| 98 | 2601.03888 | IndexTTS 2.5 Technical Report | 2026-01-07 |
| 99 | 2506.21619 | IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech | 2025-06-23 |
| 100 | 2502.05512 | IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System | 2025-02-08 |
| 101 | 2603.18090 | MOSS-TTS Technical Report | 2026-03-18 |
| 102 | 2602.10934 | MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models | 2026-02-11 |
| 103 | 2603.29339 | LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space | 2026-03-31 |
| 104 | 2603.08823 | Fish Audio S2 Technical Report | 2026-03-09 |
| 105 | 2504.18425 | Kimi-Audio Technical Report | 2025-04-25 |
| 106 | 2505.07916 | MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder | 2025-05-12 |
| — | all:"Higgs Audio" / ti:"Higgs" / all:"Boson AI" | 无相关论文条目 → Higgs Audio v2/v3 无 arXiv 论文 | — |
| — | all:"Sesame CSM conversational speech" | 无相关论文条目 → Sesame CSM 无 arXiv 论文 | — |
人类可读论文页(均返回 200): https://arxiv.org/abs/2601.15621、/2606.06928、/2509.02020、/2507.16632、/2503.01710、/2508.19205、/2512.14291、/2506.13053、/2507.09318、/2601.03888、/2506.21619、/2502.05512、/2603.18090、/2602.10934、/2603.29339、/2603.08823、/2504.18425、/2505.07916、/2412.02612、/2502.11946、/2511.15848、/2511.03601、/2509.24650
https://api.github.com/search/issues?q=...)| # | 查询 | total_count | 关键 issue |
|---|---|---|---|
| 107 | repo:OpenBMB/VoxCPM+sm_120 | 3 | #250, #326, #269 |
| 108 | repo:OpenBMB/VoxCPM+blackwell | 5 | #250, #282, #326, #102, #269 |
| 109 | repo:QwenLM/Qwen3-TTS+flash-attn | 33 | #372, #373, #350, #369, #370, #345, #179, #124 |
| 110 | repo:QwenLM/Qwen3-TTS+sm_120 | 0 | — |
| 111 | repo:microsoft/VibeVoice+5090 | 1 | #210 |
| 112 | repo:microsoft/VibeVoice+flash-attn | 24 | #377, #367, #332, #210, #185, #340, #418 |
| 113 | repo:zai-org/GLM-TTS+error | 8 | #1, #37, #3, #54, #33, #44, #24, #42 |
| 114 | repo:zai-org/GLM-TTS+blackwell+OR+5090+OR+sm_120+OR+flash | 3 | #9(RTX 5060 Ti 16G 利用率 40%), #51, #3 |
| 115 | repo:boson-ai/higgs-audio+blackwell+OR+5090+OR+sm_120+OR+flash | 5 | #39(Blackwell Docker), #177, #186, #37(TTFB 650ms), #18 |
| 116 | repo:index-tts/index-tts+blackwell+OR+5090+OR+sm_120 | 16 | #584(RTX 5090 慢), #722, #408, #717, #728, #304, #326 |
| 117 | repo:stepfun-ai/Step-Audio2+blackwell+OR+5090+OR+vllm | 34 | #57, #66, #61, #86, #68, #69 |
| 118 | repo:SesameAILabs/csm+blackwell+OR+5090+OR+sm_120 | 1 | #118(RTX 5080 sm_120 不兼容) |
| 119 | repo:k2-fsa/ZipVoice+blackwell+OR+5090+OR+sm_120 | 1 | #114(训练 CPU 内存增长,无关) |
| 120 | repo:SparkAudio/Spark-TTS+blackwell+OR+5090+OR+sm_120+OR+torch | (限流,未取得) | — |
https://github.com/OpenBMB/VoxCPM — ⚠️ 本次 curl 超时(60s);存在性由 #107~#108 的 shields 与 raw README 证实https://github.com/QwenLM/Qwen3-TTS — ⚠️ 同上https://github.com/SesameAILabs/csm — ✅ 200github.com/FireRedTeam/FireRedTTS2、stepfun-ai/Step-Audio2、SparkAudio/Spark-TTS、boson-ai/higgs-audio、microsoft/VibeVoice、zai-org/GLM-TTS、k2-fsa/ZipVoice、index-tts/index-tts、OpenMOSS/MOSS-TTS、meituan-longcat/LongCat-AudioDiT、fishaudio/fish-speech、MoonshotAI/Kimi-Audio、OpenBMB/MiniCPM-o、zai-org/GLM-4-Voice)GLM-TTS zai-org github repository、ZipVoice arxiv paper k2-fsa zero-shot TTS、IndexTTS-3 arxiv 2026、Step-Audio 2 technical report arxivFishAudio S2 open source TTS github 2026、LongCat-AudioDiT Meituan TTS github、MOSS-TTS github OpenMOSS、Qwen3-TTS-Flash open sourceQwen3-TTS-Flash open source weights or API only、Qwen3-TTS VRAM requirement GB inference 1.7Bhttps://github.com/OpenMOSS/MOSS-TTS、https://github.com/meituan-longcat/LongCat-AudioDiT、https://fish.audio/blog/fish-audio-open-sources-s2/、https://www.alibabacloud.com/help/tc/model-studio/qwen3-tts-flash| 模型 | 未查到的字段 |
|---|---|
| VoxCPM2 | TTFB(ms)、UTMOS |
| Qwen3-TTS | VRAM(GB)、RTF、UTMOS |
| FireRedTTS-2 | RTF、UTMOS、官方参数量(仅第三方表 1.5B) |
| Step-Audio 2 | TTS 侧 SeedTTS CER/SIM、UTMOS、TTFB、RTF、VRAM、官方参数量 |
| Spark-TTS | 中文官方数字、UTMOS、TTFB、VRAM |
| Higgs Audio v2 | arXiv 论文(确认不存在)、RTF、UTMOS、v2 权重的明确许可名称(HF 仅 other) |
| VibeVoice | VibeVoice-TTS 的 RTF/VRAM 官方值、UTMOS、TTS 代码(已下架) |
| Sesame CSM | arXiv 论文(确认不存在)、中文全部指标、TTFB、RTF、VRAM |
| GLM-TTS | VRAM、TTFB、RTF、UTMOS、精确参数量(HF 无 safetensors 元数据)、许可以 HF(mit) 还是 GitHub(Apache-2.0) 为准 |
| IndexTTS-2.5 | TTFB、RTF、是否/如何流式(仅社区 issue) |
| MOSS-TTS | Realtime 版参数量、各版本 VRAM 精确值、UTMOS |
| LongCat-AudioDiT | 流式支持、TTFB、RTF、VRAM |
| Fish Audio S2 Pro | test-zh SIM、UTMOS |
| ZipVoice | SeedTTS CER/SIM |
| GLM-4-Voice | 中文 SeedTTS 指标、VRAM |
| MiniCPM-o | 中文 SeedTTS 指标 |
| Kimi-Audio | License 完全未指定 |
| Qwen3-TTS-Flash | 无开源仓库/权重(闭源 API) |
| MiniMax Speech | 无开源仓库(仅论文) |
| Higgs Audio v3 | 开权重但明确非商用,未深挖参数 |
| VibeVoice-ASR-Streaming / BitNet | 本轮未深挖(非 TTS 主线) |
文档结束。本文件为原始素材(raw),后续"数字人选型结论"应在此基础上做取舍,并保留"厂商自报 vs 第三方"的口径区分。