# 第三方 TTS 评测榜单与中文实测调研（原始资料）

- **抓取日**：2026-09-22（本文所有网页均于该日实际抓取；文末「来源清单」列出全部实际抓取的 URL）
- **调研范围**：开源中文 TTS / 语音克隆选型（OmniVoice、IndexTTS-2/2.5、CosyVoice 3 / Fun-CosyVoice3-0.5B-2512、GPT-SoVITS、Fish-Speech / Fish Audio S2 Pro、MOSS-TTS、F5-TTS、MegaTTS3、Kokoro、ChatTTS、VoxCPM2、Qwen3-TTS、FireRedTTS-2、Step-Audio、Spark-TTS、Higgs Audio v2/v3、VibeVoice、GLM-TTS）
- **标注约定**（全文遵守）：
  - `[第三方]` = 与被评模型无隶属关系的独立方测得/汇编（学术论文、独立 benchmark 项目、独立博客、HF 公开投票数据）
  - `[第三方-引用]` = 第三方文章转述的某个榜单数字，我未直接访问该榜单原始页面
  - `[厂商自评]` = 模型方自己在 README / 论文里给出的对比表（含把竞品列入的表）
  - `[我方推导]` = 我从公开原始投票数据自行统计得出，**不是官方 Elo**
  - `未查到` = 本轮确实没有找到（不猜测、不编造）
- **网络限制说明**：huggingface.co 官方域名被墙；`hf-mirror.com` 可用但**明确不支持 Spaces**（返回「由于流量过大，本站暂不支持访问 Spaces 空间」）。因此所有 HuggingFace **Space** 榜单（含 TTS Arena V2）在本网络下无法直接读取；但 HF **dataset / model** 页面与 raw 文件通过镜像可正常抓取。

---

## 0. 十条最重要结论（速览）

1. **Artificial Analysis 的 TTS 榜单是 2026 年唯一"活的、有完整数字、可抓取"的第三方 Elo 榜**，但它 **只测英文（US/UK 口音）**，且 90 个模型里只有 16 个开源权重；中国开源模型（IndexTTS、CosyVoice 开源权重、GPT-SoVITS、ChatTTS、F5-TTS、MegaTTS3）**全部不在榜**。榜上最靠前的开源模型是 Breeze TTS 2（Elo 1204）、Fish Audio S2 Pro（1121）。Kokoro 82M v1.0 = **Elo 1060.98**（CI 1050–1072），是榜上最便宜（$0.65/1M 字符）。
2. **HuggingFace TTS Arena V2 的 Elo 数字本轮未能从原始页面取得**（Space 被镜像拒绝）。但找到两篇第三方文章转述其 2026 年数字：Fish Audio S2 Pro 为开源第一（ELO 1128）。
3. **Kokoro README 里提到的 "TTS Spaces Arena" 的完整投票数据是公开的**（HF dataset `Pendrokar/TTS_Arena`，2026-09-21 更新）。我自行统计 36,993 次投票事件：**Kokoro 系（hexgrad/kokoro 98.4%、hexgrad/Kokoro-API 96.7%）长期第一梯队；GPT-SoVITS-ProPlus 仅 27.5% 垫底**（标签：`[我方推导]`）。
4. **独立大规模英文 Seed-TTS-Eval 实测（VoiceHub Arena）** 给出了 33 个模型 × 1088 条英文文本的 WER/CER：**Kokoro 0.96% WER 第一，GPT-SoVITS 2.64%（第 22/33）**；该 campaign **明确声明不测 SIM/MOS**。
5. **第二个 VoiceHub campaign 提供 WavLM speaker SIM**：OmniVoice 克隆 **SIM 0.7385**、F5-TTS 克隆 **SIM 0.6691**（英文 Seed-TTS-Eval，1088 条）。
6. **GPT-SoVITS 的中文 SeedTTS 数字只有一处找到**：VoxCPM 论文（arXiv 2509.24650）在 CV3-eval 表里列 `GPT-Sovits` ZH-CER **7.34%** / EN-WER **12.5%**（明显落后于 CosyVoice2 4.08/6.32、IndexTTS2 3.58/4.45）。**该数字与论文自身 Seed-TTS-eval 表的口径不同，不可与 test-zh CER 直接对比。**
7. **MegaTTS3**：厂商论文（S-DiT，arXiv 2502.18924）自评英文 **SIM-O 0.70–0.71 / WER 1.82–1.86%**（LibriSpeech-PC）；第三方论文转引其在 Seed-TTS-eval 上 **EN-WER 2.79 / EN-SIM 77.1 / ZH-CER 1.52 / ZH-SIM 79.0**（VoxCPM 论文 Table 3；另见 RobustSpeechFlow 表）。**ChatTTS 与 Kokoro 的 test-zh/test-en WER/SIM-o 正式数字：未查到。**
8. **Kokoro 官方只支持英文**（HF 模型卡 language tag = `en`），中文需用第三方 `hexgrad/Kokoro-82M-v1.1-zh`；社区报告 KokoroSharp 版中文「口音很重、听不清」，Python 版正常。
9. **sm_120 / Blackwell**：GPT-SoVITS、CosyVoice、ChatTTS 三大框架都有 50 系显卡的真实报错记录（`no kernel image is available` / `sm_120 is not compatible with the current PyTorch installation`），根因是**整合包/requirements 里钉死的旧 PyTorch（≤2.3/2.4，只编到 sm_90）**；升级到 cu128 的 torch 2.7/2.8 可解。**索引型结论：没有一个是"sm_120 天生不支持"，都是 torch/flash-attn 版本问题。**
10. **flash-attn 在 Blackwell 上通常仍需源码编译**，但有实测可行的免编译组合（Python 3.10 + CUDA 12.8 + torch 2.7.1 + `flash_attn==2.8.0.post2 --no-build-isolation`）；Fish Audio S2 Pro 在 Blackwell 上（RTX 5090 / RTX PRO 6000）有第三方容器方案（vLLM-Omni 0.22 + torch 2.11 + cu130 + sm_120 kernel），但**上游默认配方在 sm_120 上是坏的**。

---

## 1. HuggingFace TTS Arena / TTS Spaces Arena

### 1.1 TTS Arena V2（TTS-AGI）——榜单存在但原始页面取不到

| 项目 | 结果 |
|---|---|
| Space 地址 | `https://huggingface.co/spaces/TTS-AGI/TTS-Arena-V2` |
| 本网络可读性 | ❌ 官方域名超时；`hf-mirror.com` 返回明确拒绝页（"本站暂不支持访问 Spaces 空间，可前往 Hugging Face 官网查看"） |
| 镜像可读到的唯一内容 | `https://hf-mirror.com/spaces/TTS-AGI/TTS-Arena-V2/raw/main/README.md`（176 B）内容仅为：`# TTS Arena V2` + `> Source: https://github.com/TTS-AGI/TTS-Arena` |
| 源码仓库 | `https://github.com/TTS-AGI/TTS-Arena`（可访问；Bun monorepo：`apps/web`(Next.js)、`apps/router`、`apps/docs`；Apache-2.0） |
| 文档站 | `https://docs.ttsarena.org`（可访问） |

**`docs.ttsarena.org` 明确写出的评测条件（[第三方] 榜单方自述）**：

> "TTS Arena ranks text-to-speech models by ear. You type a line, two anonymous models read it back, and you pick the one that sounds more human."
> Quick facts：**Sign in with Hugging Face to vote; accounts must be at least 30 days old. Prompts are English-only for now, capped at 1,000 characters. Models are revealed only after you vote. TTS Arena is open source under Apache 2.0.**

→ **关键限制：TTS Arena V2 的提示词目前只有英文（English-only），上限 1000 字符**，因此它**不能**用来评估中文 TTS 质量。且本轮**未查到**任何可直接抓取的 2026 年 TTS Arena V2 原始 Elo 表。

### 1.2 第三方文章转述的 TTS Arena V2 数字（[第三方-引用]）

**来源 A：Sovereign AI Blog，"Voxtral Capped at 3/10: Picking the Next Open TTS"，2026-05-12**
URL：`https://sovgrid.org/blog/strategy-tts-pivot-voxtral-ceiling/`

原文（Quick Take）：
> "TTS Arena V2 ranks **Fish Audio S2 Pro as the top open-weights model at ELO 1128**, behind five closed engines (**Realtime TTS 1.5 Max at 1208, Gemini 3.1 Flash TTS at 1206, StepAudio 2.5 TTS at 1187, ElevenLabs v3 at 1178, Inworld TTS 1 Max at 1164**). Arena measures general preference, not podcast multi-speaker dialog fitness."
> "Filtered for podcast (multi-speaker, expressivity, voice clone, **Blackwell SM12.1 compatibility**, open weights), the top three are different: **VibeVoice, Higgs Audio v2, IndexTTS-2**."

同文 "(2026-05-13 update)" 一节（作者把 TTS Arena 数字与 artificialanalysis.ai 做对照）：
> "Voxtral is on that leaderboard at **ELO 1056**, roughly tied with **Kokoro-82M v1.0**. **Fish Audio S2 Pro tops the open-weight column at 1128.** Top closed models cluster around 1180-1208…"
> "My three spike candidates are **not on the AA TTS leaderboard at all**. **VibeVoice, Higgs Audio v2, and IndexTTS-2** are either too new or have not been submitted. This is meaningful: there is no third-party benchmark to anchor the spike result against."
> "AA's leaderboard measures preference on **isolated single-sentence prompts** … It does not measure 30-second podcast monologue prosody, multi-speaker dialog turn-taking…"

> ⚠️ 注意：该文把 "TTS Arena V2" 与 "artificialanalysis.ai" 的数字混用过（1128 这个值在两个榜里都出现），因此 **1128 → Fish Audio S2 Pro** 这一条在本报告中按「第三方转述，榜单归属存疑」处理；而 1056 = Voxtral / Kokoro 这条明确标注为 **AA** 榜。

**来源 B：The AI Bench，"Best local AI voice models in 2026"，页面标注 VERIFIED SEPTEMBER 2026**
URL：`https://theaibench.ai/use-cases/voice/`

原文：
> "**TTS Arena went multi-polar in March 2026 — Fish Audio S2 Pro now leads at Elo 1128 but is non-commercial; Kokoro-82M (Apache 2.0) dropped from #1 to mid-pack on quality but remains the practical choice for English narration on CPU.**"
> "…Fish Audio S2 Pro (5B, non-commercial) if quality outranks license cleanliness."
> Verdict 推荐：**Chatterbox-Turbo (Resemble AI, MIT, Dec 15 2025)** 用于 voice cloning；**Sesame CSM-1B (Apache 2.0)** 用于 realtime conversational。

### 1.3 TTS Spaces Arena（Pendrokar，即 Kokoro README 提到的那个）——原始投票数据可抓取

这是**另一个**榜（早期由 `Pendrokar/TTS-Spaces-Arena` Space 运营），其**投票汇总数据以 HF dataset 形式完全公开**：

- Dataset：`https://hf-mirror.com/datasets/Pendrokar/TTS_Arena`，**lastModified 2026-09-21T20:39:53Z**，downloads 1913
- 文件：`tts_arena_vote_summary.tsv`、`tts_arena_vote_summary_3m.tsv`、`tts_arena_vote_summary_all.tsv`、`database.db`
- README 明确：语言 `en`，`pretty_name: TTS Spaces Arena Votes`

**TSV 结构**（我实际下载的 `tts_arena_vote_summary_all.tsv`，350,051 B）：
`spokentext | rejected（被淘汰的模型）| votes（票数）| chosen（"<数量>  <击败它的模型列表>"）| lastvote`

**`[我方推导]` 由我统计的全时段相对偏好胜率**（方法：每一行代表"在某个句子上，模型 R 被投票淘汰了 V 次，且至少有 chosen 列表中的那些模型赢过它"；因此 **wins 会被高估** —— 列表是按"曾赢过"记录的集合，不代表每次都赢；**losses 精确**。这与官方 Elo 不是一回事，只能看相对量级）。全库 36,993 次投票事件，46 个模型变体，仅列 ≥200 events：

| 胜率(推导) | events | wins | losses | 模型 |
|---|---|---|---|---|
| 98.39% | 3,726 | 3,666 | 60 | hexgrad/kokoro（原版） |
| 96.73% | 21,556 | 20,852 | 704 | hexgrad/Kokoro-API |
| 94.97% | 16,552 | 15,720 | 832 | MohamedRashad/Orpheus-TTS |
| 93.89% | 1,932 | 1,814 | 118 | Qwen/Qwen3-TTS-Voice-Design |
| 88.31% | 9,757 | 8,616 | 1,141 | coqui/xtts |
| 88.16% | 8,567 | 7,553 | 1,014 | innoai/Edge-TTS |
| 86.17% | 1,844 | 1,589 | 255 | OpenMOSS-Team/MOSS-TTS |
| 85.83% | 6,593 | 5,659 | 934 | ResembleAI/Chatterbox |
| 81.33% | 3,230 | 2,627 | 603 | fishaudio/fish-speech-1 |
| 80.73% | 5,433 | 4,386 | 1,047 | parler-tts/parler_tts |
| 80.26% | 1,768 | 1,419 | 349 | Svngoku/maskgct-audio-lab |
| 79.66% | 6,485 | 5,166 | 1,319 | Qwen/Qwen3-TTS |
| 79.40% | 3,010 | 2,390 | 620 | parler-tts/parler_tts/large |
| **79.35%** | 2,770 | 2,198 | 572 | **ByteDance/MegaTTS3** |
| 78.74% | 3,104 | 2,444 | 660 | smallest-ai-tts-lightning v3.1 |
| 77.05% | 3,111 | 2,397 | 714 | parler-tts-expresso |
| 76.01% | 792 | 602 | 190 | srinivasbilla/llasa-8b-tts |
| 75.27% | 2,386 | 1,796 | 590 | fishaudio/openaudio-s1-mini |
| 74.73% | 1,789 | 1,337 | 452 | **lj1995/GPT-SoVITS-v2** |
| 74.44% | 1,788 | 1,331 | 457 | sesame/csm-1b |
| 73.87% | 4,145 | 3,062 | 1,083 | srinivasbilla/llasa-3b-tts |
| 73.52% | 2,776 | 2,041 | 735 | Pendrokar/style-tts-2 |
| 73.17% | 1,189 | 870 | 319 | Steveeeeeeen/Zonos |
| 72.67% | 2,499 | 1,816 | 683 | **IndexTeam/IndexTTS** |
| 72.05% | 1,721 | 1,240 | 481 | Steveeeeeeen/Zonos/hybrid |
| 71.08% | 2,815 | 2,001 | 814 | **thunnai/SparkTTS** |
| 68.07% | 1,331 | 906 | 425 | OuteAI/OuteTTS-0.3-1B-Demo |
| 67.37% | 16,416 | 11,059 | 5,357 | mrfakename/E2-F5-TTS（F5-TTS） |
| 65.01% | 623 | 405 | 218 | nineninesix/KaniTTS |
| 63.59% | 4,167 | 2,650 | 1,517 | CAMB-AI/mars6-turbo-demo |
| 59.81% | 6,860 | 4,103 | 2,757 | ResembleAI/chatterbox-turbo-demo |
| 56.33% | 774 | 436 | 338 | Flux9665/EnglishToucan |
| 55.51% | 780 | 433 | 347 | nineninesix/kanitts-2-en |
| 50.10% | 1,024 | 513 | 511 | collabora/WhisperSpeech |
| 48.79% | 1,818 | 887 | 931 | Pendrokar/xVASynth-TTS |
| 48.43% | 762 | 369 | 393 | LeeSangHoon/HierSpeech_TTS |
| 47.63% | 674 | 321 | 353 | PHBJT/multi_parler_tts |
| 43.66% | 820 | 358 | 462 | Pendrokar/xVASynth-TTS/NoDeepMoji |
| 40.74% | 815 | 332 | 483 | HKUST-Audio/Llasa-1B-ft-two-speakers |
| 40.64% | 1,181 | 480 | 701 | ameerazam08/OuteTTS-0.2-500M-Demo |
| **27.54%** | 5,976 | 1,646 | 4,330 | **lj1995/GPT-SoVITS-ProPlus** |

> 解读：Kokoro 系在"单说话人 Arena"里是最强的一档（与该 Arena 的历史结论一致）；**GPT-SoVITS 的两个变体差距极大**（v2 74.7% vs ProPlus 27.5%，ProPlus 有 4330 次被淘汰，是这个 Arena 里最弱的主流模型之一）；MegaTTS3 79.4% 反而高于 F5-TTS 67.4%。这**与科技媒体上常见的排名直觉相反**，很可能因为该 Arena 是**英文单句朗读**（见下），而 ProPlus 版本是中文向微调。

**旁证（Kokoro 官方模型卡，经 Replicate 转载）**：`https://replicate.com/kjjk10/kokoro-82m/readme`
> "In the weeks leading up to its release, **Kokoro v0.19 was the #1🥇 ranked model in TTS Spaces Arena**. Kokoro achieved higher Elo in this single-voice Arena setting over other models, using fewer parameters and less data: Kokoro v0.19: 82M params, Apache, trained on <100 hours of audio; XTTS v2: 467M, CPML, >10k hours; Edge TTS: Microsoft, proprietary; MetaVoice: 1.2B, Apache, 100k hours; Parler Mini: 880M, Apache, 45k hours; Fish Speech: ~500M, CC-BY-NC-SA, 1M hours."
（注意措辞："single-voice Arena"，即 **单说话人** 榜单。这是 `[第三方-引用]/厂商模型卡转述`。）

### 1.4 ❌ 不能用的"2026 快照"：tensorfeed 聚合文件

- `https://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/<DATE>/voice-leaderboards.jsonl`
- 我抓了 5 个日期路径：`2026-05-23`、`2026-06-02`、`2026-06-18`、`2026-08-01`、`2026-09-01`、`2026-09-20` → **全部返回字节完全相同的内容（3,960 B，md5=93ac6d4f6fb5c0cf6583a9272d161c97，字段 `lastUpdated: 2026-04-30`）**。
- 其自述 `"source": "TTS Arena (Hugging Face)"`，列出 Eleven v3 (1287) → Cartesia Sonic 2 (1264) → … → **Kokoro TTS (1178, rank 8)** → Fish Audio S1 (1141, rank 10)。
- ⚠️ **结论：这是一个静态/回退文件，不能作为任何日期的 2026 年快照引用**；且其数字与 AA 榜（Kokoro 1061）和第三方转述（Fish S2 Pro 1128）都对不上。本报告**不采用**其中的 Elo 数字。

---

## 2. Artificial Analysis —— 2026 年唯一可抓取的完整第三方 TTS Elo 榜

**抓取 URL（均实际抓取，2026-09-22）**：
- `https://artificialanalysis.ai/text-to-speech`
- `https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice/open-weights` ← 主要数据来源（页面内嵌 JSON-LD 与 Next.js flight payload）
- `https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice`

**榜单口径（页面原文，`[第三方]`）**：
- 名称：**Provider Voice Arena Leaderboard** — "Compare Text to Speech (TTS) models using each provider's own native voices."
- 另有 **Controlled Voice Arena Leaderboard** — "Compare Text to Speech (TTS) models using the same 8 cloned voices (4 US, 4 UK)."
- 评分：Arena Elo，来自 **blind user votes in the Speech Arena**（用户听同一段文本的两个匿名样本，选更自然的）。FAQ 原话："Models are ranked using an Elo rating system derived from user votes in blind comparisons in the Speech Arena. Users listen to pairs of speech samples generated from the same text and choose which sounds more natural."
- **语言/口音筛选只有 `US` 和 `UK`，locale = `en`，"English (US & UK)"** → **纯英文评测，没有中文维度**。
- 口径自述："Evaluation results measured independently by Artificial Analysis"（**独立测得，非厂商自报**）。
- 榜单规模：**90 个模型，其中 16 个开源权重**（FAQ 原话）。

**页面内嵌 FAQ（JSON-LD，`[第三方]`，可直接引用）**：
> - "**Sonic 3.6** currently leads the Text to Speech Arena with an **Elo score of 1272**."
> - Top 5：1. Sonic 3.6 (1272)、2. **Qwen-Audio-3.0-TTS-Plus (1260)**、3. Realtime TTS-2 (1245)、4. Simba 3.2 (1237)、5. Luna TTS (1230)。
> - "**Kokoro 82M v1.0 is the most affordable at $0.65 per 1M characters with an Elo score of 1061.** Other affordable options include StyleTTS 2 at $2.82 per 1M characters."
> - "**Breeze TTS 2 is the highest-ranked open weights model** on the Text to Speech Leaderboard with an **Elo score of 1204**. There are **16 open weights models out of 90 total**."
> - "The top 5 open weights Text to Speech models are: **1. Breeze TTS 2 (Elo 1204), 2. Fish Audio S2 Pro (Elo 1121), 3. Step Audio EditX (Mar 2026) (Elo 1094), 4. Voxtral TTS (Elo 1075), and 5. Magpie-Multilingual 357M (Feb 2026) (Elo 1063)**."

### 2.1 全部 16 个开源权重模型（Provider Voice Arena，英文 US&UK，All categories）

数据来自页面内嵌 payload 中每个模型的默认视图行（同一模型在页面上有多个筛选切片的 Elo，我取默认视图那一条），含 95% 置信区间与投票出现次数：

| Elo | CI95 | appearances | 模型 | 权重来源 URL（页面内 `openWeightsUrl`） |
|---|---|---|---|---|
| **1204.03** | 1188–1220 | 1,374 | **Breeze TTS 2** (BreezeBlue, 2026-08-23) | `https://huggingface.co/BreezeBlue/Breeze-TTS-2` |
| **1121.03** | 1108–1134 | 2,205 | **Fish Audio S2 Pro** (Fish Audio, 2026-03-10, 4B) | `https://huggingface.co/fishaudio/s2-pro` |
| 1093.76 | — | — | **Step Audio EditX (Mar 2026)** (StepFun) | `https://huggingface.co/stepfun-ai/Step-Audio-EditX` |
| 1075.06 | 1062–1088 | 2,028 | **Voxtral TTS** (Mistral, 2026-03-26) | `https://huggingface.co/mistralai/Voxtral-4B-TTS-2603` |
| 1062.54 | — | — | **Magpie-Multilingual 357M (Feb 2026)** (NVIDIA) | `https://huggingface.co/nvidia/magpie_tts_multilingual_357m` |
| **1060.98** | 1050–1072 | 5,224 | **Kokoro 82M v1.0** (Kokoro, 2025-01-27) | `https://huggingface.co/hexgrad/Kokoro-82M` |
| 1040.56 | 1021–1061 | 1,681 | **OpenAudio S1 Mini** (Fish Audio, 2025-06-03) | `https://huggingface.co/fishaudio/s1-mini` |
| 1039.88 | — | — | **Maya1** (Maya Research) | `https://huggingface.co/maya-research/maya1` |
| **1032.15** | 1018–1046 | 1,957 | **Higgs Audio V3 TTS** (Boson AI, 2026-06-04, 4B) | `https://huggingface.co/bosonai/higgs-audio-v3-tts-4b` |
| **1020.54** | 1009–1033 | 4,595 | **Chatterbox** (Resemble AI, 2025-05-28) | `https://github.com/resemble-ai/chatterbox` |
| 1000.00 | 1000–1000 | 4,792 | **Zonos-v0.1** (Zyphra, 2025-02-10) | `https://github.com/Zyphra/Zonos` |
| **950.85** | — | — | **VibeVoice 1.5B** (Microsoft) | `https://huggingface.co/microsoft/VibeVoice-1.5B` |
| 946.34 | 932–960 | 2,959 | **OpenVoice v2** (2024-04-01) | `https://huggingface.co/myshell-ai/OpenVoiceV2` |
| 913.81 | 899–929 | 2,514 | **XTTS v2** (Coqui, 2023-11-08) | `https://huggingface.co/coqui/XTTS-v2` |
| 891.24 | 876–906 | 2,424 | **StyleTTS 2** (2023-06-13) | `https://github.com/yl4579/StyleTTS2` |
| 842.56 | — | — | **MetaVoice v1** | `https://github.com/metavoiceio/metavoice-src` |

### 2.2 相关闭源条目（同一默认视图，共 90 行中的关键项）

| rank | Elo | CI95 | appearances | 模型 | 厂商 |
|---|---|---|---|---|---|
| 0 | 1271.98 | 1255–1289 | 1,740 | Sonic 3.6 | Cartesia |
| **1** | **1260.38** | 1243–1277 | 1,438 | **Qwen-Audio-3.0-TTS-Plus** | Alibaba（页面 url 为 `/text-to-speech/model-families/cosyvoice-tts`，即归入 CosyVoice 家族；`openWeights: false`） |
| 2 | 1244.94 | 1227–1263 | 1,229 | Realtime TTS-2 | Inworld |
| 3 | 1236.88 | 1223–1251 | 2,381 | Simba 3.2 | Speechify |
| 4 | 1230.20 | 1216–1244 | 2,502 | Luna TTS | VUI Labs |
| 7 | 1198.97 | 1182–1216 | 1,262 | **StepAudio 2.5 TTS (Aug 2026)** | StepFun |
| 8 | 1198.86 | 1187–1211 | 3,415 | Gemini 3.1 Flash TTS | Google |
| 18 | 1137.90 | 1125–1151 | 2,147 | Fish Audio S2.1 Pro | Fish Audio |
| 30 | 1096.35 | 1083–1109 | 2,295 | Chatterbox HD | Resemble AI |
| 76 | 940.30 | — | — | **Qwen3 TTS Flash** | Alibaba（closed） |
| 79 | 925.98 | 913–939 | 2,854 | **Qwen3 TTS** | Alibaba（closed） |

**榜上没有的（关键空缺，`未查到`）**：IndexTTS / IndexTTS-2 / IndexTTS-2.5、CosyVoice 2/3 的**开源权重版**、GPT-SoVITS、ChatTTS、F5-TTS、MegaTTS3、Spark-TTS、GLM-TTS、FireRedTTS-2、OmniVoice、MOSS-TTS、VoxCPM2、Fish-Speech 的 v1.x 开源版。

> 注：AA 页面的 `Controlled Voice Arena`（同样 8 个克隆音色）也内嵌在同一页 payload 中，但我抓到的 payload 里同一模型存在 ~12 个不同的筛选切片 Elo（按 类别×口音 切分），**无法确定哪一条是"总榜"**，因此本报告**不引用**该板的数字，以免误标。若要引用，必须补抓该板页面并确认筛选状态。

---

## 3. 2026 年的其他第三方对比 / 基准

### 3.1 VoiceHub Arena —— 独立、大规模、可复现的英文 Seed-TTS-Eval 实测（★最有价值）

**来源**：`[第三方]` 独立研究者 kadirnar 的 VoiceHub Arena，**不隶属任何 TTS 厂商**。
- Dataset（33 模型总榜）：`https://hf-mirror.com/datasets/kadirnar/voicehub-arena-seed-tts-eval`（lastModified 2026-09-19；README + `leaderboard.csv` / `leaderboard.json` 已下载）
- Dataset（带 SIM/DNSMOS/UTMOS 的 campaign）：`https://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval`（lastModified 2026-09-19）

**协议（README 原文，条件非常明确）**：
- 测试集：**ByteDance Seed-TTS-Eval 的英文集** `en/meta.lst`，publisher revision `752f4297f090c46bb1a55a1f7439e5944ddefe8d`（"The publisher describes the English test material as originating from **Common Voice**"）
- 规模：**33 个模型家族 × 同一批 1,088 条目标文本 = 35,904 个 WAV**
- 硬件：**一台 NVIDIA A100-SXM4 40 GB**，campaign 于 **2026-09-15 完成**
- ASR：`Systran/faster-whisper-large-v3`，rev `edaa852e`，**CUDA FP16，beam 5，temperature 0，无 VAD，无 previous-text conditioning**；normalizer `whisper-normalizer==0.1.12` / `whisper_english`
- **合成 seed 42，1 次重复，identity input text；每个模型固定使用其 publisher voice/reference**（即 **fixed-voice intelligibility 评测**）
- WER/CER 为全语料 edit counts 聚合；CI 用 1,000 次 prompt-cluster bootstrap（seed 42）；RTF = 总合成时间 / 生成音频时长（不含加载/下载/ASR）；显存为**每样本峰值 CUDA allocation**
- ⚠️ **README 明确写的局限**："This is fixed-voice intelligibility evaluation, **not** an exact reproduction of the official zero-shot speaker-identity/SIM protocol. **No MOS or speaker similarity score is claimed.**"

**完整 33 模型榜（`leaderboard.csv`，按 WER 升序；我把小数换算成 %）**

| # | 模型 | checkpoint | scored | WER%↓ | CER%↓ | exact_match% | RTF | 峰值显存 MiB |
|---|---|---|---|---|---|---|---|---|
| 1 | **Kokoro** | hexgrad/Kokoro-82M | 1088 | **0.9629** | **0.2440** | 91.36 | 0.019 | 478 |
| 2 | Supertonic | Supertone/supertonic-3 | 1088 | 0.9797 | 0.2663 | 90.99 | 0.064 | 419 |
| 3 | **OmniVoice** | k2-fsa/OmniVoice | 1088 | **0.9880** | 0.2648 | 91.54 | 0.329 | 3,229 |
| 4 | SpeechT5 | microsoft/speecht5_tts | 1088 | 1.0131 | 0.2946 | 90.63 | 0.162 | 726 |
| 5 | **F5-TTS** | F5TTS_v1_Base | 1088 | **1.0550** | 0.3199 | 90.90 | 0.782 | 1,516 |
| 6 | StyleTTS2 | styletts2 epochs_2nd_00020 | 1088 | 1.0885 | 0.2991 | 90.44 | 0.042 | 1,068 |
| 7 | EchoTTS | jordand/echo-tts-base | 1088 | 1.1220 | 0.3482 | 90.35 | 0.572 | 9,363 |
| 8 | **FishTTS (S2 Pro)** | fishaudio/s2-pro | 1088 | **1.1471** | 0.3169 | 89.71 | 3.743 | 14,653 |
| 9 | XTTS | artifacts/xtts2 | 1088 | 1.1722 | 0.3928 | 89.80 | 0.377 | 2,045 |
| 10 | Zonos2 | Zyphra/ZONOS2 | 1088 | 1.1806 | 0.4211 | 90.35 | 6.228 | 15,766 |
| 11 | **Chatterbox** | ResembleAI/chatterbox | 1088 | **1.2308** | 0.3526 | 88.51 | 0.822 | 3,244 |
| 12 | **MOSS-TTS** | OpenMOSS-Team/MOSS-TTS-v1.5 | 1088 | **1.2308** | 0.4315 | 90.81 | 1.390 | 25,009 |
| 13 | OuteTTS | Llama-OuteTTS-1.0-1B | 1088 | 1.2392 | 0.4374 | 89.34 | 3.165 | 4,374 |
| 14 | **VibeVoice** | microsoft/VibeVoice-Realtime-0.5B | 1088 | **1.3146** | 0.4151 | 88.24 | 0.621 | 2,034 |
| 15 | **Qwen3-TTS** | Qwen3-TTS-12Hz-1.7B-CustomVoice | 1088 | **1.3816** | 0.4389 | 88.33 | 1.770 | 4,351 |
| 16 | NeuTTS | neuphonic/neutts-2e | 1088 | 1.4737 | 0.6204 | 88.88 | 1.704 | 3,476 |
| 17 | **CosyVoice 3**（修正后） | artifacts/cosyvoice3（=**Fun-CosyVoice3-0.5B-2512 base llm.pt**） | 1088 | **1.7416** | **0.6234** | 85.48 | 0.703 | 3,430 |
| 18 | HiggsTTS | bosonai/higgs-tts-2-3b-base | 1088 | 1.7416 | 0.6710 | 85.94 | 1.172 | 11,947 |
| 19 | InflectTTS | owensong/Inflect-Micro-v2 | 1088 | 1.7584 | 0.5892 | 85.66 | 0.014 | 174 |
| 20 | OrpheusTTS | canopylabs/orpheus-3b-0.1-ft | 1088 | 2.0012 | 0.7692 | 83.73 | 2.583 | 6,410 |
| 21 | Zonos | Zyphra/Zonos-v0.1-transformer | 1088 | 2.3026 | 0.9641 | 80.24 | 1.768 | 4,145 |
| 22 | **GPT-SoVITS** | **lj1995/GPT-SoVITS** | 1088 | **2.6375** | **0.8704** | 82.26 | 0.322 | 764 |
| 23 | MeloTTS | EN | 1088 | 2.6375 | 1.0653 | 77.67 | 0.020 | 507 |
| 24 | ParlerTTS | parler-tts-mini-v1 | 1088 | 2.6543 | 1.4462 | 80.79 | 1.877 | 4,507 |
| 25 | CSM | sesame/csm-1b | 1088 | 2.6710 | 1.4878 | 80.88 | 2.253 | 4,331 |
| 26 | **VoxCPM** | **openbmb/VoxCPM2** | 1088 | **3.3492** | 2.5992 | 83.64 | 0.874 | 5,568 |
| 27 | OpenVoice | myshell-ai/OpenVoiceV2 | 1088 | 3.6842 | 1.5905 | 72.06 | 0.042 | 753 |
| 28 | Vits | facebook/mms-tts-eng | 1088 | 6.2296 | 2.6736 | 60.85 | 0.036 | 162 |
| 29 | Bark | suno/bark-small | 1088 | 7.7367 | 4.9649 | 64.61 | 1.351 | 1,867 |
| 30 | ConversationTTS | artifacts/conversationtts | 1088 | 8.7164 | 8.2470 | 76.56 | 2.672 | 4,321 |
| 31 | Vui | vui-abraham-100m.pt | 1088 | 12.1494 | 8.3973 | 49.63 | 0.412 | 3,021 |
| 32 | Dia | nari-labs/Dia-1.6B-0626 | 1088 | 67.3533 | 59.2826 | 45.22 | 3.189 | 9,101 |
| 33 | Llasa | HKUSTAudio/Llasa-1B-Multilingual | 1088 | 73.9931 | 51.4343 | 3.13 | 0.867 | 5,498 |

**质量审计（README 原文，重要）**：
> "CosyVoice is **Fun-CosyVoice3-0.5B-2512, base llm.pt**. The archived **13.82% WER** is affected by a **confirmed HiFT implementation defect** and is excluded from ranking. The corrected full 1,088-text evaluation is verified: **WER 1.7416%, CER 0.6234%**."
> "**Llasa and Dia remain under quality review** after independent LM/codec checks."
> "Dia and Llasa have high error rates with unresolved root causes."

**同一 33 模型榜里没有的**：**ChatTTS、MegaTTS3、IndexTTS、Spark-TTS、GLM-TTS、FireRedTTS-2**（→ 这几个在该英文基准上 `未查到`）。

### 3.2 VoiceHub 的第二个 campaign：带 **WavLM speaker SIM** / DNSMOS / UTMOS22

- URL：`https://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-full-20260916/full/<method>/result.json`
- 硬件/环境（我从 `generation_provenance` 读出）：**NVIDIA RTX 6000 Ada Generation (49,140 MiB)**，`max_gpu_jobs=4`，**torch 2.8.0+cu128**，transformers 4.51.3，soundfile 0.14.0
- 指标来源：`metric_provenance` 里 ASR = faster-whisper-large-v3 (cuda/fp16/beam5)；另有 `dnsmos`、`utmos22`、`wavlm_sim`（WavLM-large ECAPA speaker SIM）
- **README 对该 campaign 的定位**："Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA speaker SIM and UTMOS22 measurements. The full campaign is still running… **Speaker SIM is N/A for generation methods without paired reference conditioning. Predicted MOS is not a human listening score.**"

**我实际抓取并聚合的 4 个 full-run（1088 条英文 Seed-TTS-Eval，全部 `evaluated=1088, generation_failures=0`）**：

| 实验 | repo | WER | CER | exact_match | DNSMOS-OVRL | UTMOS22 | **WavLM SIM** | RTF | 峰值显存 |
|---|---|---|---|---|---|---|---|---|---|
| `omnivoice--voice_clone` | k2-fsa/OmniVoice | **1.0718%** | 0.2782% | 90.81% | 3.1933 | 3.9132 | **0.7385** (n=1088, CI 0.7341–0.7428) | 0.223 | 2,207 MiB |
| `f5tts--voice_clone` | SWivid/F5-TTS | 1.2643% | 0.3853% | 88.69% | 3.0783 | 3.7026 | **0.6691** (n=1088, CI 0.6643–0.6744) | 0.606 | 781 MiB |
| `kokoro--preset_voice` | hexgrad/Kokoro-82M | **0.9462%** | 0.2410% | 91.45% | **3.4267** | **4.5037** | N/A（无参考音频） | **0.042** | 480 MiB |
| `vibevoice--preset_voice_stream` | microsoft/VibeVoice-Realtime-0.5B | 1.3481% | 0.4538% | 87.78% | 3.3381 | 4.4490 | N/A（无参考音频） | 0.629 | 2,912 MiB |

> ⚠️ 注意：同一模型在 3.1 的长表与 3.2 的 campaign 里 WER 略有差异（如 Kokoro 0.9629% vs 0.9462%，OmniVoice 0.9880% vs 1.0718%），因为**两个 campaign 是不同的运行批次/方法契约**；引用时须带 campaign 名。
> 另外 README 特别警告：`native-ada-20260916` 里的 **8 条文本 pilot 分数不得当作 full 分数**（我下载的 `native-comparison.csv` 18 行全部是 pilot，故不引用其数值）。

### 3.3 SwanBench-Speech —— 第三方学术长语音基准（浙江大學 + ByteDance）

- `https://ar5iv.labs.arxiv.org/html/2605.28618`（arXiv 2605.28618）
- `[第三方]`（学术第三方，非模型方）："SwanBench-Speech，**1,101 samples spanning 17 common speech scenarios**"，"seven metrics"，覆盖 acoustics / semantics / expressiveness 三轴，含 dialog generation，代码/演示 `https://swanaigc.github.io/#bench`
- 该文指出"existing test scenarios are often confined to limited domains"、"current models still struggle in highly expressive scenarios and exhibit a notable gap in consistency and hierarchy compared to real recordings"
- ⚠️ 我抓到的 HTML 中未提取到逐模型数值表（表格为图/LaTeX 化）；**具体模型排名：未查到**。

### 3.4 中文第三方实测/文章（含可信度评价）

**(a) 博客园 sensorsen《语音模型 2026 开源 TTS 选型指南：六款主流语音合成模型实测对比》（2026）** `[第三方]`（个人博客，非严谨实测）
URL：`https://www.cnblogs.com/sensorsen/p/21367537`
- 作者自述仍是"开源TTS本地部署和运行示例"，且开头写："**还有 omnivoice，稍后验证后上效果**" → 覆盖 6 款：CosyVoice 2、Qwen3-TTS、Fish Speech 1.5、IndexTTS2、F5-TTS、Spark-TTS
- 作者实测/主观结论（原文）：
  - **Qwen3-TTS**："Qwen3-TTS 语音模型**不能加减速**，实际应用效果远不及 cosyvoice。**非常不推荐**。"（同时承认其"端到端延迟仅 97ms"）
  - **IndexTTS2**："在语音克隆的情感表达上做到了目前最优"，"7 种基础情感独立控制、Token 级时长精确控制（**误差 < 0.02%**）"、"中英双语，**55K 小时**训练数据（30K 中文 + 25K 英文）"、"**FP16 下约 7.8GB 显存**，RTX 3060 8GB 即可运行"
  - **CosyVoice 2**："单路实时推理约 **4GB 显存**，RTX 4090 可跑 2 路并发"，"3 秒零样本克隆 + 流式推理（首包延迟约 **1.5 秒**）"
  - **Fish Speech 1.5**：基础推理 ≥4GB、克隆 ≥6GB；"务必使用 Python 3.12"（3.13/3.14 依赖冲突）、"切到 v1.5.0 或 v1.5.1 tag，不要用 main 分支"
  - **Spark-TTS**："显存需求仅 **1-2GB**"，"**16kHz 采样率偏低**，长文本表现一般"
  - **F5-TTS**：MIT，~3GB
- ⚠️ 该文"效果对比（客观指标）"表（号称"按论文数据和社区 Benchmark 汇总"）：CosyVoice 2 英文 WER 2.07% / 中文 WER 2.43% / 情感相似度 0.831；IndexTTS2 1.88% / 2.12% / 0.872；Spark-TTS 2.43% / 2.87% / 0.847；Fish Speech 1.5 3.50% / 1.30%。**这些数字与官方 SeedTTS 口径（CosyVoice2 test-zh CER 1.45 / test-en WER 2.57）对不上，来源不可追溯 → 本报告标记为「低可信，勿引用」。**
- 该文末尾贴的其实是 CosyVoice README 的官方表（见 4.3）。

**(b) CSDN《和GPT-SoVITS比如何？两款热门中文TTS横向对比》（2026-05-28）** `[第三方，低可信/AI 生成痕迹明显]`
URL：`https://blog.csdn.net/weixin_42602241/article/details/156922929`
- 对比 **GPT-SoVITS vs "IndexTTS2 V23（科哥团队维护）"**（注意：**"科哥版"不是 bilibili 官方 index-tts**）
- 唯一量化项（原文表格）：**推理延迟 RTF：GPT-SoVITS ~0.8–1.2（RTX 3060）；IndexTTS2 V23 ~0.6–0.9（同设备）**；**显存：GPT-SoVITS ≥6GB(FP16)；IndexTTS2 V23 ≥4GB(FP16)**；模型体积 ~5GB vs ~3.2GB
- 原文结论："若追求极致音色拟合与跨语言能力，GPT-SoVITS 更具优势；若侧重快速部署、情感可控性和低资源消耗，IndexTTS2 表现更优。"
- ⚠️ 文章含大量模板化表述，且把 IndexTTS2 归给"科哥团队"，**型号归属有误**，数字建议仅作旁证。

**(c) 博客园 kacoro《tts哪家强？》（2025-07-10）** `[第三方]`（真实使用日志）
URL：`https://www.cnblogs.com/kacoro/p/18977125`
- 线上用 minimax + GPT-SOVITS；开源试过 index-tts / CosyVoice / gpt-sovits；作者结论"最终确定的是 SoVITS"
- 对比表：index-tts 样本需求"10s 内？"、不支持自训练；CosyVoice2 30s、支持；GPT-SoVITS 3~10s、支持自训练
- 真实体感（原文）："对于带有方言，果然还是得自己训练。那些只考 3~10 秒就想出比较满意的效果。是很难的。"
- GPT-SoVITS 训练用 1 小时素材 + 12 轮 SoVITS/25 轮 GPT 后："训练出来的音色与语气已经很接近了。**但是存在大舌头的情况**"；调参经验："**合成轮数越低口齿会更清晰，但语气会差一些**"、"高轮次语气好，但存在**声音泄露**问题"、"关闭推理后再重新打开是解决声音泄露很好的办法"
- ⚠️ 文中那张"数据时长/轮次/最小显存/MOS"表，作者自己标注："**上面这份表，是 deepseek 给我的，也不知道准不准确**" → **不可引用其 MOS**。

**(d) 腾讯云开发者社区《我们把 Kokoro 1.1 量化到 FP16，并用它替换了浏览器里的中文 Piper 配音》（2026-08-08）** `[第三方]`
URL：`https://cloud.tencent.cn/developer/article/2722884`
- 摘要仅为："在最新发布的 **Timeline Studio v1.0.0** 中，我们完成了一次重要的**中文 AI 配音升级**"；标签 `#Kokoro1.1 #文字转语音 #onnx`
- ⚠️ 正文主体在我的抓取下未展开（页面为前端渲染/含登录门槛），**具体中文质量数据未查到**；仅能证明"有团队在 2026 年用 Kokoro 1.1 (ONNX/FP16) 做生产级中文配音"。

**(e) SiliconFlow《终极指南 — 2026年最佳开源 Text-to-Speech 模型》** `[厂商内容营销]`
URL：`https://www.siliconflow.com/zh/articles/best-open-source-text-to-speech-models`
- 原文："我们针对 2026 年推荐的前三款模型是 **Fish Speech V1.5、CosyVoice2-0.5B 和 IndexTTS-2**"
- 自称"我们与行业业内人士合作，测试了关键基准上的性能" → 但未给出可核对的数字表，**应视为厂商营销，不是第三方基准**
- 另见其姊妹篇 `https://www.siliconflow.com/zh/articles/best-open-source-models-for-voice-cloning`（本轮未抓取正文）

**(f) 中文 Arena 的公开数据（未被充分利用的线索）** `[第三方原始数据]`
- `https://hf-mirror.com/datasets/JacobLinCool/zh-tw-tts-arena-votes`（lastModified 2026-07-24，license cc-by-4.0）
  README 原文："Append-only log of blind A/B preference votes collected by the zh-tw-tts-arena Space. Schema: `ts, session, sentence_id, model_a, model_b, condition_a, condition_b, winner(a|b|tie), dwell_s, ua, app_version`."（4 个 jsonl 文件）
  → **这是本轮唯一找到的「中文（繁體）盲测 Arena」原始投票数据**；我只抓了 README，**未统计 → 中文 Arena 排名：未查到（可后续补充）**

### 3.5 英文第三方对比文章

**(a) The AI Bench — "Best local AI voice models in 2026"（VERIFIED SEPTEMBER 2026）** `[第三方]`
URL：`https://theaibench.ai/use-cases/voice/`
- 见 1.2 来源 B 引文；另有分层推荐：Frontier (64+GB) → Qwen3-Omni-30B-A3B FP16（omnimodal）、Fish Audio S2 Pro（5B，非商用）；STT → Canary-Qwen 2.5B（"#1 on Open ASR Leaderboard at 5.63% WER"）+ WhisperX

**(b) Sovereign AI Blog（2026-05-12）** `[第三方]` — 见 1.2 来源 A；另有重要的**工程现实**结论：
- Voxtral open-checkpoint **两个不重叠的失败模式**："Turns under 100 chars produce filler hallucinations, turns over 350 chars flatten into staccato. No turn-length sweet spot reaches release quality. **Capped at 3/10 best case**"；"The `instructions` parameter is **silently ignored**. The `ref_audio` parameter **crashes** the engine because the encoder weights stayed gated in Mistral's hosted product. **No speed knob exists**."
- 该文的生产环境：**DGX Spark / GB10 Blackwell**（ARM v9.2-A + **SM12.1**）+ CUDA 13 wheels；作者把"Blackwell SM12.1 兼容性"作为**筛选硬条件**
- 选型结论：需要 multi-speaker + 表达力 + 克隆 + 速度控制 + 开源许可 → **只 spike 三个：VibeVoice（社区 fork）、Higgs Audio v2、IndexTTS-2**；并指出 **IndexTTS-2 明显值得一试**（该文是独立第三方，非 IndexTTS 方）
- 该文提到 VibeVoice："Microsoft's original release describes it as 'designed for up to **4 distinct speakers in a single generation of up to 90 minutes**' using next-token diffusion. It was an **ICLR 2026 oral**."；"**Microsoft pulled the original repository in September 2025** over concerns about deepfake misuse"

**(c) Pinggy / Replicate 等其他英文页**：`https://pinggy.io/blog/best_open_source_self_hosted_text_to_speech_models/`（本轮**未抓取正文**）；`https://replicate.com/kjjk10/kokoro-82m/readme`（已抓，见 1.3）

---

## 4. SeedTTS test-zh / test-en 的 WER/CER 与 SIM-o / SIM-r 数字

> **先给结论（针对指定的 4 个老模型）**：
> - **GPT-SoVITS**：Seed-TTS test-zh/test-en 的 WER/SIM **未查到**；只查到 **CV3-eval 上的 ZH-CER 7.34% / EN-WER 12.5%**（VoxCPM 论文，见 4.1）；另有 VoiceHub 的**英文** WER 2.6375%/CER 0.8704%（见 3.1）。
> - **MegaTTS3**：厂商论文给的是 **LibriSpeech-PC 的 SIM-O / WER（英文，非 SeedTTS）**；SeedTTS 口径的 **EN-WER 2.79 / EN-SIM 77.1 / ZH-CER 1.52 / ZH-SIM 79.0** 来自**他人论文的转引表**（见 4.1、4.2）。
> - **Kokoro**：SeedTTS 的 WER/CER/SIM-o **未查到**（Kokoro 无官方论文）；最接近的是 VoiceHub 的英文 Seed-TTS-Eval WER **0.9629%**（见 3.1）与 KaniTTS 的英文 MOS/WER/CER（见 4.4）。
> - **ChatTTS**：SeedTTS 的 WER/CER/SIM **完全未查到**（本轮在所有抓取的论文表格中都没有出现 ChatTTS）。

### 4.1 VoxCPM 论文（OpenBMB）—— 含 MegaTTS3 与 GPT-Sovits 的两张关键表 `[厂商自评]`（论文作者是 VoxCPM 方）

URL：`https://arxiv.org/html/2509.24650v1`（arXiv 2509.24650，标题 *VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning*）

**Table 3: Performance on Seed-TTS-eval Benchmark**（原文表头：`Model | Params | Open-Source | EN WER↓ SIM↑ | ZH CER↓ SIM↑ | Hard CER↓ SIM↑`）
所有 SIM 为百分数（0-100），与其它论文的 0-1 口径不同，注意换算：

| Model | Params | OS | EN-WER↓ | EN-SIM↑ | ZH-CER↓ | ZH-SIM↑ | Hard-CER↓ | Hard-SIM↑ |
|---|---|---|---|---|---|---|---|---|
| **MegaTTS3** (Jiang et al., 2025) | 0.5B | ✗ | **2.79** | **77.1** | **1.52** | **79.0** | – | – |
| DiTAR | 0.6B | ✗ | 1.69 | 73.5 | 1.02 | 75.3 | – | – |
| CosyVoice3 (0.5B) | 0.5B | ✗ | 2.02 | 71.8 | 1.16 | 78.0 | 6.08 | 75.8 |
| CosyVoice3 (1.5B) | 1.5B | ✗ | 2.22 | 72.0 | 1.12 | 78.1 | 5.83 | 75.8 |
| Seed-TTS | – | ✗ | 2.25 | 76.2 | 1.12 | 79.6 | 7.59 | 77.6 |
| MiniMax-Speech | – | ✗ | 1.65 | 69.2 | 0.83 | 78.3 | – | – |
| F5-TTS | 0.3B | ✓ | 2.00 | 67.0 | 1.53 | 76.0 | 8.67 | 71.3 |
| MaskGCT | – | ✓ | 2.62 | 71.7 | 2.27 | 77.4 | – | – |
| CosyVoice | 0.3B | ✓ | 4.29 | 60.9 | 3.63 | 72.3 | 11.75 | 70.9 |
| CosyVoice2 | 0.5B | ✓ | 3.09 | 65.9 | 1.38 | 75.7 | 6.83 | 72.4 |
| SparkTTS | 0.5B | ✓ | 3.14 | 57.3 | 1.54 | 66.0 | – | – |
| FireRedTTS | 0.5B | ✓ | 3.82 | 46.0 | 1.51 | 63.5 | 17.45 | 62.1 |
| FireRedTTS-2 | – | ✓ | 1.95 | 66.5 | 1.14 | 73.6 | – | – |
| Qwen2.5-Omni | 7B | ✓ | 2.72 | 63.2 | 1.70 | 75.2 | 7.97 | 74.7 |
| OpenAudio-s1-mini | 0.5B | ✓ | 1.94 | 55.0 | 1.18 | 68.5 | 23.37 | 64.3 |
| IndexTTS 2 | 1.5B | ✓ | 2.23 | 70.6 | 1.03 | 76.5 | 7.12 | 75.5 |
| VibeVoice | 1.5B | ✓ | 3.04 | 68.9 | 1.16 | 74.4 | – | – |
| HiggsAudio-v2 | 3B | ✓ | 2.44 | 67.7 | 1.50 | 74.0 | 55.07 | 65.6 |
| VoxCPM-Emilia | 0.5B | ✓ | 2.34 | 68.1 | 1.11 | 74.0 | 12.46 | 69.8 |
| VoxCPM | 0.5B | ✓ | 1.85 | 72.9 | 0.93 | 77.2 | 8.87 | 73.0 |

**Table 4: Performance on CV3-eval Benchmark**（`*denotes close-sourced systems`）——**这是唯一找到的 GPT-Sovits 数字**：

| Model | CV3-EVAL ZH-CER↓ | CV3-EVAL EN-WER↓ | CV3-Hard-ZH CER↓ | SIM↑ | DNSMOS↑ | CV3-Hard-EN WER↓ | SIM↑ | DNSMOS↑ |
|---|---|---|---|---|---|---|---|---|
| F5-TTS | 5.47 | 8.90 | – | – | – | – | – | – |
| SparkTTS | 5.15 | 11.0 | – | – | – | – | – | – |
| **GPT-Sovits** | **7.34** | **12.5** | – | – | – | – | – | – |
| CosyVoice2 | 4.08 | 6.32 | 12.58 | 72.6 | 3.81 | 11.96 | 66.7 | 3.95 |
| OpenAudio-s1-mini | 4.00 | 5.54 | 18.1 | 58.2 | 3.77 | 12.4 | 55.7 | 3.89 |
| IndexTTS2 | 3.58 | 4.45 | 12.8 | 74.6 | 3.65 | 8.78 | 74.5 | 3.80 |
| HiggsAudio-v2 | 9.54 | 7.89 | 41.0 | 60.2 | 3.39 | 10.3 | 61.8 | 3.68 |
| CosyVoice3-0.5B* | 3.89 | 5.24 | 14.15 | 78.6 | 3.75 | 9.04 | 75.9 | 3.92 |
| CosyVoice3-1.5B* | 3.91 | 4.99 | 9.77 | 78.5 | 3.79 | 10.55 | 76.1 | 3.95 |
| VoxCPM-Emilia | 4.47 | 5.23 | 22.2 | 62.6 | 3.47 | 10.00 | 62.6 | 3.68 |
| VoxCPM | 3.40 | 4.04 | 12.9 | 66.1 | 3.59 | 7.89 | 64.3 | 3.74 |

> **口径警告**：CV3-eval 的 CER 数值（4~8%）**远高于** Seed-TTS-eval test-zh 的 CER（1~2%），两套测试集的正文长度/难度不同。**不能用 CV3-eval 的 7.34% 去和 test-zh 的 1.5% 对比**。

### 4.2 MegaTTS3 官方论文（S-DiT）—— 厂商自评，英文 LibriSpeech-PC `[厂商自评]`

URL：`https://arxiv.org/html/2502.18924v1`（标题 *Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis*，即 MegaTTS3 / S-DiT）

**主表（Table 5，英文 LibriSpeech-PC 类测试集）**：
| Model | #Params | Training Data | SIM-O↑ | SIM-R↑ | WER↓ | CMOS↑ | SMOS↑ | RTF↓ |
|---|---|---|---|---|---|---|---|---|
| GT | – | – | 0.68 | – | 1.94% | +0.12 | 3.92 | – |
| VALL-E 2* | 0.4B | LibriHeavy | 0.64 | 0.68 | 2.44% | – | – | – |
| VoiceBox† | 0.4B | Collected (60kh) | 0.64 | 0.67 | 2.03% | −0.20 | 3.81 | 0.340 |
| DiTTo-TTS* | 0.7B | Collected (55kh) | 0.62 | 0.65 | 2.56% | – | – | – |
| NaturalSpeech 3† | 0.5B | LibriLight | 0.67 | 0.76 | 1.81% | −0.10 | 3.95 | 0.296 |
| CosyVoice | 0.4B | Collected (172kh) | 0.62 | – | 2.24% | −0.18 | 3.93 | 1.375 |
| MaskGCT | 1.0B | Emilia (100kh) | 0.69 | – | 2.63% | – | – | – |
| F5-TTS | 0.3B | Emilia (100kh) | 0.66 | – | 1.96% | −0.12 | 3.96 | 0.307 |
| **S-DiT（=MegaTTS3）** | 0.3B | LibriLight | **0.71** | **0.78** | **1.82%** | 0.00 | 3.98 | 0.188 |
| S-DiT-accelerated | 0.3B | LibriLight | 0.70 | 0.78 | 1.86% | −0.03 | 3.96 | 0.124 |

其余表（同一论文）：Table 6 SIM-O 0.70 / WER 2.31%；Table 11 `Ours 0.71 / 1.82%`；Table 12 规模消融（0.5B: SIM-O 0.66 / WER 2.10%；1.5B: 0.72 / 1.98%；7.0B: 0.74 / 1.90%）；Table 14 长文本 WER 2.39%（vs CosyVoice 5.52%、VoiceCraft 12.81%）。

### 4.3 第三方论文转引的 MegaTTS3（SeedTTS 口径）

**(a) RobustSpeechFlow（arXiv 2605.22083）的 Table 7** `[第三方学术]`
URL：`https://arxiv.org/html/2605.22083v1`（*RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching*）
> 表头：`Model | Params | WER↓ | SIM↑`
> **MegaTTS3 [23] | 0.5B | 2.79 | 0.77**；Seed-TTS DiT [21] | – | 1.73 | 0.79；DiTAR [5] | 0.6B | 1.69 | 0.74；MiniMax-Speech [24] | – | 1.65 | 0.69；F5-TTS [2] | 0.3B | 2.00 | 0.67；CosyVoice3 [25] | 1.5B | 2.22 | 0.72；Spark-TTS [26] | 0.5B | 3.14 | 0.57；OpenAudio S1-Mini [27] | 0.5B | 1.94 | 0.55；IndexTTS2 [28] | 1.5B | 2.23 | 0.71；VibeVoice [7] | 1.5B | 3.04 | 0.69；VoxCPM-Emilia [6] | 0.5B | 2.34 | 0.68；VoxCPM [6] | 0.5B | 1.85 | 0.73

**(b) WavTTS（arXiv 2606.03455）Table 10 —— Seed-TTS test-en / test-zh** `[第三方学术]`
URL：`https://arxiv.org/html/2606.03455v1`（*WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling*）
表头：`Model | Params | Data(hrs) | Seed-TTS test-en: WER%↓ SIM-o↑ UTMOS↑ | Seed-TTS test-zh: CER%↓ SIM-o↑ UTMOS↑`

| Model | Params | Data | en WER% | en SIM-o | en UTMOS | zh CER% | zh SIM-o | zh UTMOS |
|---|---|---|---|---|---|---|---|---|
| Ground Truth | – | – | 1.79 | 0.73 | 3.53 | 1.25 | 0.75 | 2.78 |
| CosyVoice | 416M | 170K Multi. | 4.29 | 0.61 | – | 3.63 | 0.72 | – |
| CosyVoice 2 | 618M | 167K Multi. | 2.57 | 0.65 | – | 1.45 | 0.75 | – |
| Llasa-1B | 1370M | 250K Multi. | 3.22 | 0.57 | – | 1.89 | 0.67 | – |
| Spark-TTS | 507M | 102K Multi. | 1.98 | 0.58 | – | 1.20 | 0.67 | – |
| MaskGCT | 1048M | 100K Emilia | 2.36 | 0.71 | 3.57 | 2.48 | 0.77 | 2.64 |
| E2-TTS | 333M | 100K Emilia | 2.21 | 0.71 | 3.20 | 1.97 | 0.73 | 2.27 |
| F5-TTS | 336M | 100K Emilia | 1.65 | 0.66 | 3.73 | 1.55 | 0.75 | 2.94 |
| ZipVoice | 123M | 100K Emilia | 1.60 | 0.70 | 3.83 | 1.40 | 0.75 | 3.15 |
| LongCat-AudioDiT | 1420M | 100K Multi. | 1.94 | 0.76 | 3.80 | 1.10 | 0.81 | 3.16 |
| WavTTS | 673M | 100K Emilia | 1.50 | 0.65 | 3.92 | 1.59 | 0.73 | 3.08 |
（**无 GPT-SoVITS / MegaTTS3 / Kokoro / ChatTTS**）

**(c) CaT-TTS（arXiv 2509.22062）Table 11** `[第三方学术]`
URL：`https://arxiv.org/html/2509.22062v1`（*Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling*）
表头 `Model | test-zh WER%↓ SIM↑ | test-en WER%↓ SIM↑ | test-hard WER%↓ SIM↑`：

| Model | zh WER% | zh SIM | en WER% | en SIM | hard WER% | hard SIM |
|---|---|---|---|---|---|---|
| MaskGCT | 2.27 | 0.774 | 2.62 | 0.714 | 10.27 | 0.748 |
| E2 TTS (32 NFE) | 1.97 | 0.730 | 2.19 | 0.710 | – | – |
| F5-TTS (32 NFE) | 1.56 | 0.741 | 1.83 | 0.647 | 8.67 | 0.713 |
| Seed-TTS | 1.12 | 0.796 | 2.25 | 0.762 | 7.59 | 0.776 |
| FireRedTTS | 1.51 | 0.635 | 3.82 | 0.460 | 17.45 | 0.621 |
| CosyVoice | 3.63 | 0.723 | 4.29 | 0.609 | 11.75 | 0.709 |
| CosyVoice 2 | 1.45 | 0.748 | 2.57 | 0.652 | 6.83 | 0.724 |
| CosyVoice 3-0.5B | 1.16 | 0.780 | 2.02 | 0.718 | 6.08 | 0.758 |
| QTTS | 1.66 | 0.648 | 3.17 | 0.652 | 14.45 | 0.641 |
| Spark-TTS | 1.20 | 0.672 | 1.98 | 0.584 | – | – |
| Llasa-1B/3B/8B-250k | 1.89/1.60/1.59 | 0.668/0.675/0.684 | 3.22/3.14/2.97 | 0.572/0.579/0.574 | 12.13/13.37/11.09 | 0.638/0.652/0.660 |
| CaT-TTS | 1.56 | 0.678 | 2.35 | 0.668 | 9.75 | 0.674 |
（**无 GPT-SoVITS / MegaTTS3 / Kokoro / ChatTTS**）

### 4.4 Kokoro 专项

- **官方支持语言**：`https://hf-mirror.com/api/models/hexgrad/Kokoro-82M` 返回 tags 含 `"text-to-speech","en"` → **官方仅英文**。
- **中文/多语第三方权重确实存在**（`https://hf-mirror.com/api/models?search=Kokoro`，按下载量排序）：
  - `hexgrad/Kokoro-82M-v1.1-zh`（downloads 22,373）
  - `onnx-community/Kokoro-82M-v1.1-zh-ONNX`（tags 含 `en, zh`）
  - `aufklarer/Kokoro-82M-CoreML`（tags 含 `en, zh`）
  - 其他语种：`contextboxai/Kokoro-Vietnamese`、`zaakirio/kokoro-ru`、`Thorsten-Voice/Kokoro`(de)、`TilLabs/kokoro-tts-kazakh` 等
- **中文质量第三方报告**：
  - `https://github.com/Lyrcaxis/KokoroSharp/issues/5`（2025-02-11，issue 正文）："I tried this library to generate Chinese speech, but I found that it has a **very heavy accent, and it's a bit hard to hear what's being said** in the generated content, but using the **python version of the kokoro library, the generated Chinese speech files don't have this problem**."（即问题出在 C# 端实现，而非模型本身）
- **KaniTTS 的 Kokoro 基准（第三方数据集，英文）** `[第三方]`
  - `https://hf-mirror.com/datasets/KaniTTS-research-team/kokoro_banchmark`（lastModified 2026-03-18；1 split，3600 行，7.2 MB parquet，含 `MOS/NOI/DIS/COL/LOUD/UTMOS/WER/CER`）
  - 我下载 parquet 并**逐 model 求均值**（样本文本为英文，如 "The birch canoe slid on the smooth planks."）：

| model（Kokoro 英文音色） | n | MOS | UTMOS | WER | CER |
|---|---|---|---|---|---|
| kokoro_82m_en_am_fenrir | 720 | 4.986 | 3.760 | 0.013 | 0.005 |
| kokoro_82m_en_af_heart | 720 | 4.981 | 3.759 | 0.013 | 0.004 |
| kokoro_82m_en_am_michael | 720 | 4.923 | 3.814 | 0.012 | 0.004 |
| kokoro_82m_en_af_bella | 720 | 4.538 | 3.122 | 0.010 | 0.004 |
| kokoro_82m_en_bf_emma | 720 | 4.453 | 3.501 | 0.012 | 0.004 |
（MOS ≈ 4.45–4.99 偏高，疑似非严格人工 MOS；WER/CER 为比例值，×100 即百分数：1.0–1.3% WER、0.4–0.5% CER。**注意这是英文**。）

### 4.5 其它厂商自评表（不满足"第三方"要求，但含关键竞品横排，供交叉验证）

**(a) CosyVoice README 官方表（Alibaba FunAudioLLM）** `[厂商自评]`
URL：`https://raw.githubusercontent.com/FunAudioLLM/CosyVoice/main/README.md`
表头：`Model | Open-Source | Size | test-zh CER%↓ SS%↑ | test-en WER%↓ SS%↑ | test-hard CER%↓ SS%↑`

| Model | OS | Size | test-zh CER | test-zh SS | test-en WER | test-en SS | test-hard CER | test-hard SS |
|---|---|---|---|---|---|---|---|---|
| Human | – | – | 1.26 | 75.5 | 2.14 | 73.4 | – | – |
| Seed-TTS | ❌ | – | 1.12 | 79.6 | 2.25 | 76.2 | 7.59 | 77.6 |
| MiniMax-Speech | ❌ | – | 0.83 | 78.3 | 1.65 | 69.2 | – | – |
| F5-TTS | ✅ | 0.3B | 1.52 | 74.1 | 2.00 | 64.7 | 8.67 | 71.3 |
| Spark TTS | ✅ | 0.5B | 1.2 | 66.0 | 1.98 | 57.3 | – | – |
| CosyVoice2 | ✅ | 0.5B | 1.45 | 75.7 | 2.57 | 65.9 | 6.83 | 72.4 |
| FireRedTTS2 | ✅ | 1.5B | 1.14 | 73.2 | 1.95 | 66.5 | – | – |
| Index-TTS2 | ✅ | 1.5B | 1.03 | 76.5 | 2.23 | 70.6 | 7.12 | 75.5 |
| VibeVoice-1.5B | ✅ | 1.5B | 1.16 | 74.4 | 3.04 | 68.9 | – | – |
| VibeVoice-Realtime | ✅ | 0.5B | – | – | 2.05 | 63.3 | – | – |
| HiggsAudio-v2 | ✅ | 3B | 1.50 | 74.0 | 2.44 | 67.7 | – | – |
| VoxCPM | ✅ | 0.5B | 0.93 | 77.2 | 1.85 | 72.9 | 8.87 | 73.0 |
| GLM-TTS | ✅ | 1.5B | 1.03 | 76.1 | – | – | – | – |
| GLM-TTS RL | ✅ | 1.5B | 0.89 | 76.4 | – | – | – | – |
| **Fun-CosyVoice3-0.5B-2512** | ✅ | 0.5B | **1.21** | **78.0** | **2.24** | **71.8** | **6.71** | **75.8** |
| Fun-CosyVoice3-0.5B-2512_RL | ✅ | 0.5B | **0.81** | 77.4 | **1.68** | 69.5 | **5.44** | 75.0 |
（**无 GPT-SoVITS / MegaTTS3 / Kokoro / ChatTTS**）

**(b) IndexTTS 官方 README（bilibili）** `[厂商自评]`
URL：`https://raw.githubusercontent.com/index-tts/index-tts/main/README.md`（同表亦可从 `https://raw.githubusercontent.com/T8mars/indextts25-desktop-t8/main/README.md` 得到同内容并含 Table 2）
> "**Table 1: Zero-shot TTS on CV3-Eval** (Arabic uses an in-house test set). **†Cited from the original paper.**"

| Model | Params | zh WER% | zh SS% | en WER% | en SS% | es WER% | es SS% | ja WER% | ja SS% | ar WER% | ar SS% | Avg WER% | Avg SS% |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| **VoxCPM2** | 2B | 3.88 | 74.99 | 5.13 | 71.57 | 5.49 | 74.67 | 6.69 | 72.90 | 14.94 | 65.99 | 7.22 | 72.02 |
| **OmniVoice** | 0.8B | **3.41** | 72.99 | **3.62** | 70.13 | 3.52 | 74.14 | 5.38 | 70.49 | 17.88 | 64.22 | 6.76 | 70.39 |
| Moss-TTS 1.5 | 8B | 4.02 | 72.68 | 4.45 | 67.46 | 3.83 | 71.75 | 10.97 | 68.71 | 23.71 | 62.21 | 9.40 | 68.56 |
| **CosyVoice3-0.5B** | 0.5B | 3.84 | **80.01** | 4.88 | 74.16 | 4.04 | 78.85 | – | 76.36 | – | – | – | – |
| CosyVoice3-1.5B | 1.5B | 3.91† | – | 4.99† | – | 4.47† | – | 7.57† | – | – | – | – | – |
| FireRedTTS-2 | 1.5B | 8.22 | 68.10 | 14.92 | 56.93 | – | – | – | – | – | – | – | – |
| Fish Audio S2 Pro | 4B | 3.62 | 67.79 | 3.83 | 61.66 | 2.93 | 67.44 | 5.15 | 66.15 | 14.15 | 59.43 | 5.94 | 64.49 |
| Qwen3-TTS | 1.7B | 3.27 | 73.02 | 5.06 | 67.17 | 2.87 | 73.17 | 5.89 | 70.18 | – | – | – | – |
| IndexTTS2.5 | 0.8B | 4.36 | 77.10 | 5.12 | 68.06 | 3.75 | 76.39 | 5.66 | 74.62 | 14.88 | 69.74 | 6.75 | 73.18 |
| **IndexTTS2.5-RL** | 0.8B | 3.93 | 77.92 | 3.89 | 67.79 | 3.33 | 76.68 | 5.30 | 75.41 | 13.58 | 70.36 | **6.00** | **73.63** |

Table 2（跨语言，中文 prompt → 目标语言）关键行：OmniVoice zh→en WER 3.74 / SS 64.91；zh→es 5.84 / 62.08；zh→ja 9.09 / 69.06；zh→ar 19.80 / 65.27；平均 9.62 / 65.33。IndexTTS2.5-RL 平均 6.17 / 70.20。
（注：**这套表是 IndexTTS 团队做的**，把 OmniVoice / VoxCPM2 / CosyVoice3 / Fish S2 Pro / Qwen3-TTS 都排了进去 → **属于厂商对比，不是第三方中立评测**；其中带 † 的来自原论文，其余应为该团队自测。）

**(c) VoxCPM 官方 README（含 zh 版）** `[厂商自评]` — URL：`https://raw.githubusercontent.com/OpenBMB/VoxCPM/main/README.md`
提供 **Seed-TTS-eval 表**（VoxCPM2 2B：zh 1.84 / SIM 75.3；en 0.97 / SIM 79.5；hard 8.13 / 75.3；FishAudio S2 4B：zh **0.99**、en **0.54**；LongCat-Audio-DiT 3.5B：zh 1.50 / **SIM 78.6**、en 1.09 / **81.8**；Qwen3-TTS 1.7B：zh 1.23、en 1.22；MOSS-TTS：zh 1.85 / SIM 73.4、en 1.20 / 78.8）
及 **CV3-eval 多语表**（VoxCPM2: zh 3.65 / en 5.00 / hard-zh 8.55 / hard-en 8.48；Fish Audio S2: 2.65 / 2.43 / 9.10 / 4.40）
及 **MiniMax-Multilingual-Test**（Chinese: FishAudio S2 **0.730**、Qwen3-TTS 0.928、VoxCPM2 1.136、MiniMax 2.252、ElevenLabs 16.026）

**(d) OmniVoice README（k2-fsa）** `[厂商自评]` — URL：`https://raw.githubusercontent.com/k2-fsa/OmniVoice/main/README.md`
- 无 SeedTTS WER/SIM 表；给的是 **FlashInfer 加速基准**：
  > "Benchmark (**seed-tts zh testset, 2020 samples / 3.3h audio, voice cloning, single H100, fp16, num_step=32**; Average RTF as reported by `omnivoice-infer-batch`, outputs ASR-verified lossless)"
  | batch | baseline RTF | FlashInfer RTF | speedup |
  |---|---|---|---|
  | 1 | 0.0899 | 0.0430 | 2.1× |
  | 1 + CUDA graph | – | 0.0367 | 2.4× |
  | 2 | 0.0480 | 0.0245 | 2.0× |
  | 4 | 0.0331 | 0.0152 | 2.2× |
  | 8 | 0.0298 | **0.0115** | **2.6×** |
- 另：voice design 仅中英训练；"**3–10 秒**参考音频"建议；"For better results with Arabic numerals, normalize them to words first"；引文 `arXiv:2604.00688`（*OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models*）

**(e) Fish Speech README（fishaudio）** `[厂商自评]` — URL：`https://raw.githubusercontent.com/fishaudio/fish-speech/main/README.md`
- **许可证重要变化**：代码与权重改为 **FISH AUDIO RESEARCH LICENSE**（不再是 Apache-2.0）："This codebase and its associated model weights are released under **[FISH AUDIO RESEARCH LICENSE]**. We will take action against any violation of the license."
- 部署路径指向 SGLang-Omni / vLLM-Omni 配方（`https://github.com/vllm-project/vllm-omni/blob/main/recipes/fishaudio/Fish-Speech-S2-Pro.md`）
- 该 README 片段中**未含 SeedTTS WER/SIM 表**（本轮抓取的正文部分）

---

## 5. sm_120 / Blackwell / RTX 50 系（含 5060 Ti 16GB）兼容性实况

### 5.1 结论速查表（全部有 URL 证据）

| 框架 | 是否需 flash-attn | sm_120 现成 wheel？ | 已知报错（原文） | 可行 workaround | 证据 URL |
|---|---|---|---|---|---|
| **GPT-SoVITS** | 否（核心路径不用） | 是（用 cu128 官方 torch wheel 即可） | `NVIDIA GeForce RTX 5070 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation. The current PyTorch install supports CUDA capabilities sm_37 sm_50 sm_60 sm_61 sm_70 sm_75 sm_80 sm_86 sm_90 compute_37`；`RuntimeError: "no kernel image is available" on new NVIDIA GPUs (e.g., RTX 5090 / sm_120)`；5060 Ti 上 `s2_train.py` 的 `mp.spawn` 报错 | **换 cu128 的 torch（≥2.7）**；整合包自带旧 torch 是根因。另有 PR #2774「Fix s1_train DDP crash on Windows single-GPU (sm_120 / Blackwell)」 | 见 5.2 |
| **CosyVoice / Fun-CosyVoice3** | 否（官方路径用 SDPA；有 triton-TRTLLM 可选） | 是 | Error A: `'torch.distributed_c10d.ProcessGroup' object has no attribute 'options'`（torch 2.4+ 移除）；Error B（按 requirements 降到 torch 2.3.1 时）：`NVIDIA GeForce RTX 5060 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation (supports up to sm_90). Note: RTX 50-series cards require modern CUDA toolkits and PyTorch versions not covered by the current requirements.`；Error C: `ERROR: Cannot install hyperpyyaml==1.2.2 and ruamel.yaml<=0.17.21 because these package versions have conflicting dependencies` | **升级到 torch 2.7/2.8+cu128，不要按老 requirements 降级**；ruamel.yaml 冲突需手工解 | `https://github.com/QwenAudio/CosyVoice/issues/1815` |
| **ChatTTS（jianchang512/ChatTTS-ui 整合包）** | 否 | 是（但**整合包内置旧 torch**） | `NVIDIA GeForce RTX 5070 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation…`（整合包换 50 系卡后运行 app.exe 即报） | 换掉整合包内 torch 为 cu128 版 | `https://github.com/jianchang512/ChatTTS-ui/issues/290` |
| **Fish-Speech / Fish Audio S2 Pro** | **是**（上游配方在 sm_120 会坏） | 需用预置容器 | "the standard way to add S2-Pro's DAC codec to it **breaks on new RTX 5090 / RTX PRO 6000 Blackwell cards (FlashAttention-3 ships kernels only for Hopper sm90a / Blackwell-Ultra sm120a, and SGLang is blocked on sm_120)**" | 第三方项目 `Genesis1231/fish-s2-rtx`：容器 pinned 到 **vLLM-Omni 0.22 的 torch 2.11 + cu130 + sm_120 kernels**；另注"Prefix caching is intentionally off — it intermittently produces empty audio in vLLM-Omni 0.22" | `https://github.com/Genesis1231/fish-s2-rtx` |
| **index-tts / IndexTTS2** | 未查到确切结论 | – | 有 issue「能用RTX50系显卡跑吗？」（#242，**本轮抓取超时，正文未取得 → 未查到**）；另有 `torch.AcceleratorError: CUDA error: operation not permitted when stream is capturing` (#677) | 未查到 | issue 列表来自 GitHub 搜索 API 结果：`https://github.com/index-tts/index-tts/issues/242`（正文未获取） |
| **flash-attn（通用）** | — | **有可用的免编译组合** | "many people struggle with building (or installing) flash attention" | **实测可行**：`Python=3.10 + CUDA=12.8 + torch 2.7.1`（`pip install torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1 --index-url https://download.pytorch.org/whl/cu128`）然后 `pip install flash_attn==2.8.0.post2 torch==2.7.1 --no-build-isolation` | `https://github.com/Dao-AILab/flash-attention/issues/2016` |
| **F5-TTS** | – | – | **未查到**（`sneekes.app` 的 "F5-TTS Installation Guide for RTX 5070 on WSL2" 两次抓取均超时，正文未取得） | 未查到 | 线索 URL（**未成功抓取，不得引用其内容**）：`https://sneekes.app/posts/f5-tts-installation-guide-for-rtx-5070-on-wsl2/` |
| **MegaTTS3 / MOSS-TTS** | – | – | **未查到** | 未查到 | — |

### 5.2 GPT-SoVITS 细节（逐个 issue）

- **`https://github.com/RVC-Boss/GPT-SoVITS/issues/2205`**「50系列显卡不支持」（opened 2025-03-19，Closed）
  训练时报：`NVIDIA GeForce RTX 5070 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation. The current PyTorch install supports CUDA capabilities sm_37 sm_50 sm_60 sm_61 sm_70 sm_75 sm_80 sm_86 sm_90 compute_37. If you want to use the NVIDIA GeForce RTX 5070 Ti GPU with PyTorch, please check the instructions at https://pytorch.org/get-started/locally/`
- **`https://github.com/RVC-Boss/GPT-SoVITS/issues/2514`**「[Bug Report] RuntimeError: "no kernel image is available" on new NVIDIA GPUs (e.g., RTX 5090 / sm_120)」（2025-07-09，状态 In follow-up）
  报告人自述："it seems the current **PyTorch binaries do not yet support the architecture** of the latest GPUs… I've attempted to fix this by upgrading PyTorch to the latest available versions, including the **nightly builds**, but it appears that even the newest pre-release binaries do not yet include support for sm_120. The command I used was: `pip install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/cu121`"
  → ⚠️ 注意：该用户装的是 **cu121 nightly**，而 sm_120 需要 **cu128 及以上**，所以他的尝试方式本身就是错的（这条记录很好用来说明"为什么很多人以为不支持"）。
- **`https://github.com/RVC-Boss/GPT-SoVITS/issues/2192`**「NVIDIA RTX 5070 with CUDA 12.8 (sm_120), Error Message」（Closed）
  在 `GPT_SoVITS/inference_webui.py` 阶段出现 transformers/pytree 的 FutureWarning 之后报错（正文含完整 traceback）。
- **`https://github.com/RVC-Boss/GPT-SoVITS/issues/2393`**「搭载 5060 ti显卡，sovits模型训练错误」/ **#2394**「搭载5060 ti显卡上 训练sovits模型出现的报错。gpt模型是可以训练的」（均 Closed）
  traceback 指向 `GPT_SOVITS/s2_train.py` 第 601 行 `main()` → 第 57 行 `mp.spawn(...)` → `torch/multiprocessing/spawn.py`。
  → **关键信息：5060 Ti 上「GPT 模型可以训练，SoVITS 模型训练报错」**（#2394 标题直述）。
- 相关 PR：**`https://github.com/RVC-Boss/GPT-SoVITS/pull/2774`**「Fix s1_train DDP crash on Windows single-GPU (sm_120 / Blackwell)」（open）——说明 **DDP 在 sm_120 上仍有坑**。

### 5.3 其他 Blackwell 相关实况

- **Sovereign AI Blog（2026-05）** 记录了在 **DGX Spark / GB10（ARM v9.2-A + Blackwell SM12.1 + CUDA 13 wheels）** 上跑开源 TTS 的一系列真实故障：
  - "**Voxtral Stage 1 OOM on GB10: Why `--enforce-eager` Is Not Enough**"（2026-04-25，容器起不来）
  - "the **3.5-hour deadlock** that was really an `AttributeError`（Python init-order bug）that masqueraded as a **Blackwell GPU hang**"（2026-05-03）
  - "the **three-line vllm-omni patch for Blackwell** on the same day, the upstream fix that **made Voxtral actually run on SM12.1**"
  - 明确把 **"ARM v9.2-A and Blackwell SM12.1 compatibility. The Spark uses CUDA 13 wheels for PyTorch; most TTS engines ship for x86 CUDA 12. Bring-up risk is real and engine-specific."** 列为选型硬指标
  - 该文提到 `microsoft/VibeVoice-Realtime-0.5B` 有"verified setup thread for DGX Spark with CUDA 13 and aarch64"，是候选里**唯一有公开 Spark 部署记录**的引擎
  - URL：`https://sovgrid.org/blog/strategy-tts-pivot-voxtral-ceiling/`
- **ggml/CUDA 的相关 bug（旁证 sm_120 支持仍在补齐）**：`https://github.com/ggml-org/ggml/issues/1466`「ggml-cuda: flash-attn MMA picker has **no Blackwell (sm_120) entry** — silently uses Ampere config」（**本报告仅通过搜索发现该 URL，未抓取正文 → 内容未核实**）

---

## 6. 明确"未查到"的项（供后续补做）

1. **TTS Arena V2 的原始 Elo 榜单页**（HF Space 被镜像拒绝；本网络无法读取）——只有第三方转述（Fish Audio S2 Pro 1128）。
2. **Chinese / 中文 Arena 的正式排名**——只有 `JacobLinCool/zh-tw-tts-arena-votes` 的原始 jsonl（未统计）。
3. **ChatTTS 的任何 SeedTTS WER/CER/SIM 数字**（本轮所有抓取的论文表、benchmark 表中均无 ChatTTS）。
4. **Kokoro 的 test-zh CER / SIM-o**（Kokoro 无官方论文；第三方 benchmark 只覆盖英文）。
5. **GPT-SoVITS 的 SeedTTS test-zh CER / test-en WER 正式数字**（只有 CV3-eval 口径的 7.34 / 12.5）。
6. **MegaTTS3 的 SeedTTS test-zh SIM-o 原始出处**（只有 VoxCPM 论文与 RobustSpeechFlow 的转引）。
7. **F5-TTS 在 RTX 5070 上的安装实录**（sneekes.app 两次抓取超时）。
8. **index-tts issue #242「能用RTX50系显卡跑吗？」正文**（两次抓取超时）→ IndexTTS 在 sm_120 上的实测结论缺口。
9. **MOSS-TTS / MegaTTS3 / GLM-TTS 的 sm_120 记录**。
10. **SwanBench-Speech 的逐模型分数表**（HTML 中表格未以可解析文本呈现）。

---

## 7. 来源清单（本文实际抓取过的全部 URL）

### 7.1 TTS Arena / 榜单
1. `https://hf-mirror.com/spaces/TTS-AGI/TTS-Arena-V2`（镜像拒绝页）
2. `https://hf-mirror.com/spaces/TTS-AGI/TTS-Arena-V2/raw/main/README.md`（176 B）
3. `https://huggingface.co/spaces/TTS-AGI/TTS-Arena-V2`（官方域名，超时不可达）
4. `https://ttsarena.org`、`https://ttsarena.org/leaderboard`、`https://ttsarena.org/api/leaderboard`（均为静态 TTS-AGI 项目页）
5. `https://docs.ttsarena.org`
6. `https://github.com/TTS-AGI/TTS-Arena`
7. `https://raw.githubusercontent.com/TTS-AGI/TTS-Arena/main/README.md`
8. `https://artificialanalysis.ai/text-to-speech`
9. `https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice/open-weights`
10. `https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice`
11. `https://artificialanalysis.ai/api/text-to-speech/leaderboard`（404）
12. `https://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-05-23/voice-leaderboards.jsonl`
13. `https://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-06-02/voice-leaderboards.jsonl`
14. `https://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-06-18/voice-leaderboards.jsonl`
15. `https://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-08-01/voice-leaderboards.jsonl`
16. `https://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-09-01/voice-leaderboards.jsonl`
17. `https://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-09-20/voice-leaderboards.jsonl`
18. `https://hf-mirror.com/datasets/Pendrokar/TTS_Arena/raw/main/README.md`
19. `https://hf-mirror.com/datasets/Pendrokar/TTS_Arena/raw/main/tts_arena_vote_summary_all.tsv`
20. `https://hf-mirror.com/api/datasets/Pendrokar/TTS_Arena`
21. `https://replicate.com/kjjk10/kokoro-82m/readme`

### 7.2 独立 benchmark 数据集（HF，经镜像）
22. `https://hf-mirror.com/datasets/kadirnar/voicehub-arena-seed-tts-eval/raw/main/README.md`
23. `https://hf-mirror.com/datasets/kadirnar/voicehub-arena-seed-tts-eval/raw/main/leaderboard.csv`
24. `https://hf-mirror.com/api/datasets/kadirnar/voicehub-arena-seed-tts-eval`
25. `https://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/README.md`
26. `https://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-20260916/native-comparison.csv`
27. `https://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-full-20260916/full/omnivoice--voice_clone/result.json`
28. `https://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-full-20260916/full/f5tts--voice_clone/result.json`
29. `https://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-full-20260916/full/kokoro--preset_voice/result.json`
30. `https://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-full-20260916/full/vibevoice--preset_voice_stream/result.json`
31. `https://hf-mirror.com/api/datasets/VoiceHub/voicehub-arena-seed-tts-eval`
32. `https://hf-mirror.com/datasets/KaniTTS-research-team/kokoro_banchmark/raw/main/README.md`
33. `https://hf-mirror.com/api/datasets/KaniTTS-research-team/kokoro_banchmark`
34. `https://hf-mirror.com/datasets/KaniTTS-research-team/kokoro_banchmark/resolve/main/data/train-00000-of-00001.parquet`
35. `https://hf-mirror.com/datasets/JacobLinCool/zh-tw-tts-arena-votes/raw/main/README.md`
36. `https://hf-mirror.com/api/datasets?search=tts+arena&limit=30`、`https://hf-mirror.com/api/datasets?search=TTS-AGI&limit=30`
37. `https://hf-mirror.com/api/models/hexgrad/Kokoro-82M`
38. `https://hf-mirror.com/api/models?search=Kokoro&limit=40&sort=downloads&direction=-1`
39. `https://hf-mirror.com/api/datasets/VoiceHub/voicehub-arena-seed-tts-eval`（文件清单）
40. `https://hf-mirror.com/api/datasets/kadirnar/voicehub-arena-seed-tts-eval`（文件清单）

### 7.3 论文 / README（原始文件）
41. `https://arxiv.org/html/2509.24650v1`（VoxCPM；含 Table 3 Seed-TTS-eval、Table 4 CV3-eval 及 **GPT-Sovits**）
42. `https://arxiv.org/html/2502.18924v1`（MegaTTS3 / S-DiT 官方论文，英文 LibriSpeech-PC SIM-O/WER）
43. `https://arxiv.org/html/2509.22062v1`（CaT-TTS，SeedTTS test-zh/en/hard）
44. `https://arxiv.org/html/2605.22083v1`（RobustSpeechFlow，含 MegaTTS3 转引）
45. `https://arxiv.org/html/2606.03455v1`（WavTTS，Seed-TTS test-en/zh SIM-o）
46. `https://ar5iv.labs.arxiv.org/html/2605.28618`（SwanBench-Speech 长语音基准）
47. `https://raw.githubusercontent.com/FunAudioLLM/CosyVoice/main/README.md`
48. `https://raw.githubusercontent.com/k2-fsa/OmniVoice/main/README.md`
49. `https://raw.githubusercontent.com/fishaudio/fish-speech/main/README.md`
50. `https://raw.githubusercontent.com/index-tts/index-tts/main/README.md`
51. `https://raw.githubusercontent.com/T8mars/indextts25-desktop-t8/main/README.md`（IndexTTS 官方 Table 1/Table 2 全文）
52. `https://raw.githubusercontent.com/OpenBMB/VoxCPM/main/README.md`
53. `https://github.com/T8mars/indextts25-desktop-t8`
54. `https://github.com/FunAudioLLM/CV3-Eval`（页面已抓，本报告未引用其内容）

### 7.4 sm_120 / Blackwell
55. `https://github.com/RVC-Boss/GPT-SoVITS/issues/2514`
56. `https://github.com/RVC-Boss/GPT-SoVITS/issues/2192`
57. `https://github.com/RVC-Boss/GPT-SoVITS/issues/2205`
58. `https://github.com/RVC-Boss/GPT-SoVITS/issues/2393`
59. `https://github.com/RVC-Boss/GPT-SoVITS/pull/2774`（经 GitHub 搜索 API 结果发现；**正文未单独抓取**）
60. `https://github.com/RVC-Boss/GPT-SoVITS/issues/2626`（经搜索发现；正文未抓取）
61. `https://github.com/QwenAudio/CosyVoice/issues/1815`
62. `https://github.com/jianchang512/ChatTTS-ui/issues/290`
63. `https://github.com/Dao-AILab/flash-attention/issues/2016`
64. `https://github.com/Genesis1231/fish-s2-rtx`
65. `https://api.github.com/search/issues?q=repo:RVC-Boss/GPT-SoVITS+sm_120`（搜索 API，仅用于发现；后触发速率限制）
66. `https://api.github.com/search/issues?q=repo:RVC-Boss/GPT-SoVITS+5060`
67. `https://api.github.com/search/issues?q=repo:index-tts/index-tts+5060`
68. `https://huggingface.co/2Noise/ChatTTS/discussions/38`（**仅搜索结果中出现，未抓取**）

### 7.5 第三方文章 / 中文来源
69. `https://sovgrid.org/blog/strategy-tts-pivot-voxtral-ceiling/`
70. `https://theaibench.ai/use-cases/voice/`
71. `https://www.cnblogs.com/sensorsen/p/21367537`
72. `https://blog.csdn.net/weixin_42602241/article/details/156922929`
73. `https://www.cnblogs.com/kacoro/p/18977125`
74. `https://www.siliconflow.com/zh/articles/best-open-source-text-to-speech-models`
75. `https://cloud.tencent.cn/developer/article/2722884`
76. `https://github.com/Lyrcaxis/KokoroSharp/issues/5`
77. `https://pinggy.io/blog/best_open_source_self_hosted_text_to_speech_models/`（**未抓取正文**）
78. `https://sneekes.app/posts/f5-tts-installation-guide-for-rtx-5070-on-wsl2/`（**两次抓取均超时，内容未取得**）
79. `https://huggingface.co/spaces/Pendrokar/TTS-Spaces-Arena`（README 中引用的原 Space 地址；本网络不可达）

### 7.6 辅助
- `web_search` 工具（用于发现候选 URL，共 6 轮）：包括 "TTS Arena leaderboard Elo 2026"、"2026 开源 TTS 对比 实测"、"GPT-SoVITS ChatTTS MegaTTS3 test-zh CER SIM-o comparison table paper"、"CosyVoice RTX 5090 sm_120 flash-attn issue github" 等查询。
- `http://export.arxiv.org/api/query`（arXiv API，用于检索候选论文）
- 本地工具：`/tmp/dh/f.sh`（curl 抓取）、`/tmp/dh/txt.sh`（HTML→文本）、自写 `tab.py`（arXiv 表格抽取）、venv + pyarrow（读 KaniTTS parquet）

---

*本文件为原始资料汇编，未做结论提炼；所有数字均标注了来源与"厂商自评/第三方"属性，未标注者请在引用前回溯到 7 节对应 URL 二次核对。*
