[第三方] = 与被评模型无隶属关系的独立方测得/汇编(学术论文、独立 benchmark 项目、独立博客、HF 公开投票数据)[第三方-引用] = 第三方文章转述的某个榜单数字,我未直接访问该榜单原始页面[厂商自评] = 模型方自己在 README / 论文里给出的对比表(含把竞品列入的表)[我方推导] = 我从公开原始投票数据自行统计得出,不是官方 Elo未查到 = 本轮确实没有找到(不猜测、不编造)hf-mirror.com 可用但明确不支持 Spaces(返回「由于流量过大,本站暂不支持访问 Spaces 空间」)。因此所有 HuggingFace Space 榜单(含 TTS Arena V2)在本网络下无法直接读取;但 HF dataset / model 页面与 raw 文件通过镜像可正常抓取。Pendrokar/TTS_Arena,2026-09-21 更新)。我自行统计 36,993 次投票事件:Kokoro 系(hexgrad/kokoro 98.4%、hexgrad/Kokoro-API 96.7%)长期第一梯队;GPT-SoVITS-ProPlus 仅 27.5% 垫底(标签:[我方推导])。GPT-Sovits ZH-CER 7.34% / EN-WER 12.5%(明显落后于 CosyVoice2 4.08/6.32、IndexTTS2 3.58/4.45)。该数字与论文自身 Seed-TTS-eval 表的口径不同,不可与 test-zh CER 直接对比。en),中文需用第三方 hexgrad/Kokoro-82M-v1.1-zh;社区报告 KokoroSharp 版中文「口音很重、听不清」,Python 版正常。no kernel image is available / sm_120 is not compatible with the current PyTorch installation),根因是整合包/requirements 里钉死的旧 PyTorch(≤2.3/2.4,只编到 sm_90);升级到 cu128 的 torch 2.7/2.8 可解。索引型结论:没有一个是"sm_120 天生不支持",都是 torch/flash-attn 版本问题。flash_attn==2.8.0.post2 --no-build-isolation);Fish Audio S2 Pro 在 Blackwell 上(RTX 5090 / RTX PRO 6000)有第三方容器方案(vLLM-Omni 0.22 + torch 2.11 + cu130 + sm_120 kernel),但上游默认配方在 sm_120 上是坏的。| 项目 | 结果 |
|---|---|
| Space 地址 | https://huggingface.co/spaces/TTS-AGI/TTS-Arena-V2 |
| 本网络可读性 | ❌ 官方域名超时;hf-mirror.com 返回明确拒绝页("本站暂不支持访问 Spaces 空间,可前往 Hugging Face 官网查看") |
| 镜像可读到的唯一内容 | https://hf-mirror.com/spaces/TTS-AGI/TTS-Arena-V2/raw/main/README.md(176 B)内容仅为:# TTS Arena V2 + > Source: https://github.com/TTS-AGI/TTS-Arena |
| 源码仓库 | https://github.com/TTS-AGI/TTS-Arena(可访问;Bun monorepo:apps/web(Next.js)、apps/router、apps/docs;Apache-2.0) |
| 文档站 | https://docs.ttsarena.org(可访问) |
docs.ttsarena.org 明确写出的评测条件([第三方] 榜单方自述):
"TTS Arena ranks text-to-speech models by ear. You type a line, two anonymous models read it back, and you pick the one that sounds more human." Quick facts:Sign in with Hugging Face to vote; accounts must be at least 30 days old. Prompts are English-only for now, capped at 1,000 characters. Models are revealed only after you vote. TTS Arena is open source under Apache 2.0.
→ 关键限制:TTS Arena V2 的提示词目前只有英文(English-only),上限 1000 字符,因此它不能用来评估中文 TTS 质量。且本轮未查到任何可直接抓取的 2026 年 TTS Arena V2 原始 Elo 表。
来源 A:Sovereign AI Blog,"Voxtral Capped at 3/10: Picking the Next Open TTS",2026-05-12 URL:https://sovgrid.org/blog/strategy-tts-pivot-voxtral-ceiling/
原文(Quick Take):
"TTS Arena V2 ranks Fish Audio S2 Pro as the top open-weights model at ELO 1128, behind five closed engines (Realtime TTS 1.5 Max at 1208, Gemini 3.1 Flash TTS at 1206, StepAudio 2.5 TTS at 1187, ElevenLabs v3 at 1178, Inworld TTS 1 Max at 1164). Arena measures general preference, not podcast multi-speaker dialog fitness." "Filtered for podcast (multi-speaker, expressivity, voice clone, Blackwell SM12.1 compatibility, open weights), the top three are different: VibeVoice, Higgs Audio v2, IndexTTS-2."
同文 "(2026-05-13 update)" 一节(作者把 TTS Arena 数字与 artificialanalysis.ai 做对照):
"Voxtral is on that leaderboard at ELO 1056, roughly tied with Kokoro-82M v1.0. Fish Audio S2 Pro tops the open-weight column at 1128. Top closed models cluster around 1180-1208…" "My three spike candidates are not on the AA TTS leaderboard at all. VibeVoice, Higgs Audio v2, and IndexTTS-2 are either too new or have not been submitted. This is meaningful: there is no third-party benchmark to anchor the spike result against." "AA's leaderboard measures preference on isolated single-sentence prompts … It does not measure 30-second podcast monologue prosody, multi-speaker dialog turn-taking…"
⚠️ 注意:该文把 "TTS Arena V2" 与 "artificialanalysis.ai" 的数字混用过(1128 这个值在两个榜里都出现),因此 1128 → Fish Audio S2 Pro 这一条在本报告中按「第三方转述,榜单归属存疑」处理;而 1056 = Voxtral / Kokoro 这条明确标注为 AA 榜。
来源 B:The AI Bench,"Best local AI voice models in 2026",页面标注 VERIFIED SEPTEMBER 2026 URL:https://theaibench.ai/use-cases/voice/
原文:
"TTS Arena went multi-polar in March 2026 — Fish Audio S2 Pro now leads at Elo 1128 but is non-commercial; Kokoro-82M (Apache 2.0) dropped from #1 to mid-pack on quality but remains the practical choice for English narration on CPU." "…Fish Audio S2 Pro (5B, non-commercial) if quality outranks license cleanliness." Verdict 推荐:Chatterbox-Turbo (Resemble AI, MIT, Dec 15 2025) 用于 voice cloning;Sesame CSM-1B (Apache 2.0) 用于 realtime conversational。
这是另一个榜(早期由 Pendrokar/TTS-Spaces-Arena Space 运营),其投票汇总数据以 HF dataset 形式完全公开:
https://hf-mirror.com/datasets/Pendrokar/TTS_Arena,lastModified 2026-09-21T20:39:53Z,downloads 1913tts_arena_vote_summary.tsv、tts_arena_vote_summary_3m.tsv、tts_arena_vote_summary_all.tsv、database.dben,pretty_name: TTS Spaces Arena VotesTSV 结构(我实际下载的 tts_arena_vote_summary_all.tsv,350,051 B): spokentext | rejected(被淘汰的模型)| votes(票数)| chosen("<数量> <击败它的模型列表>")| lastvote
[我方推导] 由我统计的全时段相对偏好胜率(方法:每一行代表"在某个句子上,模型 R 被投票淘汰了 V 次,且至少有 chosen 列表中的那些模型赢过它";因此 wins 会被高估 —— 列表是按"曾赢过"记录的集合,不代表每次都赢;losses 精确。这与官方 Elo 不是一回事,只能看相对量级)。全库 36,993 次投票事件,46 个模型变体,仅列 ≥200 events:
| 胜率(推导) | events | wins | losses | 模型 |
|---|---|---|---|---|
| 98.39% | 3,726 | 3,666 | 60 | hexgrad/kokoro(原版) |
| 96.73% | 21,556 | 20,852 | 704 | hexgrad/Kokoro-API |
| 94.97% | 16,552 | 15,720 | 832 | MohamedRashad/Orpheus-TTS |
| 93.89% | 1,932 | 1,814 | 118 | Qwen/Qwen3-TTS-Voice-Design |
| 88.31% | 9,757 | 8,616 | 1,141 | coqui/xtts |
| 88.16% | 8,567 | 7,553 | 1,014 | innoai/Edge-TTS |
| 86.17% | 1,844 | 1,589 | 255 | OpenMOSS-Team/MOSS-TTS |
| 85.83% | 6,593 | 5,659 | 934 | ResembleAI/Chatterbox |
| 81.33% | 3,230 | 2,627 | 603 | fishaudio/fish-speech-1 |
| 80.73% | 5,433 | 4,386 | 1,047 | parler-tts/parler_tts |
| 80.26% | 1,768 | 1,419 | 349 | Svngoku/maskgct-audio-lab |
| 79.66% | 6,485 | 5,166 | 1,319 | Qwen/Qwen3-TTS |
| 79.40% | 3,010 | 2,390 | 620 | parler-tts/parler_tts/large |
| 79.35% | 2,770 | 2,198 | 572 | ByteDance/MegaTTS3 |
| 78.74% | 3,104 | 2,444 | 660 | smallest-ai-tts-lightning v3.1 |
| 77.05% | 3,111 | 2,397 | 714 | parler-tts-expresso |
| 76.01% | 792 | 602 | 190 | srinivasbilla/llasa-8b-tts |
| 75.27% | 2,386 | 1,796 | 590 | fishaudio/openaudio-s1-mini |
| 74.73% | 1,789 | 1,337 | 452 | lj1995/GPT-SoVITS-v2 |
| 74.44% | 1,788 | 1,331 | 457 | sesame/csm-1b |
| 73.87% | 4,145 | 3,062 | 1,083 | srinivasbilla/llasa-3b-tts |
| 73.52% | 2,776 | 2,041 | 735 | Pendrokar/style-tts-2 |
| 73.17% | 1,189 | 870 | 319 | Steveeeeeeen/Zonos |
| 72.67% | 2,499 | 1,816 | 683 | IndexTeam/IndexTTS |
| 72.05% | 1,721 | 1,240 | 481 | Steveeeeeeen/Zonos/hybrid |
| 71.08% | 2,815 | 2,001 | 814 | thunnai/SparkTTS |
| 68.07% | 1,331 | 906 | 425 | OuteAI/OuteTTS-0.3-1B-Demo |
| 67.37% | 16,416 | 11,059 | 5,357 | mrfakename/E2-F5-TTS(F5-TTS) |
| 65.01% | 623 | 405 | 218 | nineninesix/KaniTTS |
| 63.59% | 4,167 | 2,650 | 1,517 | CAMB-AI/mars6-turbo-demo |
| 59.81% | 6,860 | 4,103 | 2,757 | ResembleAI/chatterbox-turbo-demo |
| 56.33% | 774 | 436 | 338 | Flux9665/EnglishToucan |
| 55.51% | 780 | 433 | 347 | nineninesix/kanitts-2-en |
| 50.10% | 1,024 | 513 | 511 | collabora/WhisperSpeech |
| 48.79% | 1,818 | 887 | 931 | Pendrokar/xVASynth-TTS |
| 48.43% | 762 | 369 | 393 | LeeSangHoon/HierSpeech_TTS |
| 47.63% | 674 | 321 | 353 | PHBJT/multi_parler_tts |
| 43.66% | 820 | 358 | 462 | Pendrokar/xVASynth-TTS/NoDeepMoji |
| 40.74% | 815 | 332 | 483 | HKUST-Audio/Llasa-1B-ft-two-speakers |
| 40.64% | 1,181 | 480 | 701 | ameerazam08/OuteTTS-0.2-500M-Demo |
| 27.54% | 5,976 | 1,646 | 4,330 | lj1995/GPT-SoVITS-ProPlus |
解读:Kokoro 系在"单说话人 Arena"里是最强的一档(与该 Arena 的历史结论一致);GPT-SoVITS 的两个变体差距极大(v2 74.7% vs ProPlus 27.5%,ProPlus 有 4330 次被淘汰,是这个 Arena 里最弱的主流模型之一);MegaTTS3 79.4% 反而高于 F5-TTS 67.4%。这与科技媒体上常见的排名直觉相反,很可能因为该 Arena 是英文单句朗读(见下),而 ProPlus 版本是中文向微调。
旁证(Kokoro 官方模型卡,经 Replicate 转载):https://replicate.com/kjjk10/kokoro-82m/readme
"In the weeks leading up to its release, Kokoro v0.19 was the #1🥇 ranked model in TTS Spaces Arena. Kokoro achieved higher Elo in this single-voice Arena setting over other models, using fewer parameters and less data: Kokoro v0.19: 82M params, Apache, trained on <100 hours of audio; XTTS v2: 467M, CPML, >10k hours; Edge TTS: Microsoft, proprietary; MetaVoice: 1.2B, Apache, 100k hours; Parler Mini: 880M, Apache, 45k hours; Fish Speech: ~500M, CC-BY-NC-SA, 1M hours."
(注意措辞:"single-voice Arena",即 单说话人 榜单。这是 [第三方-引用]/厂商模型卡转述。)
https://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/<DATE>/voice-leaderboards.jsonl2026-05-23、2026-06-02、2026-06-18、2026-08-01、2026-09-01、2026-09-20 → 全部返回字节完全相同的内容(3,960 B,md5=93ac6d4f6fb5c0cf6583a9272d161c97,字段 lastUpdated: 2026-04-30)。"source": "TTS Arena (Hugging Face)",列出 Eleven v3 (1287) → Cartesia Sonic 2 (1264) → … → Kokoro TTS (1178, rank 8) → Fish Audio S1 (1141, rank 10)。抓取 URL(均实际抓取,2026-09-22):
https://artificialanalysis.ai/text-to-speechhttps://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice/open-weights ← 主要数据来源(页面内嵌 JSON-LD 与 Next.js flight payload)https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice榜单口径(页面原文,[第三方]):
US 和 UK,locale = en,"English (US & UK)" → 纯英文评测,没有中文维度。页面内嵌 FAQ(JSON-LD,[第三方],可直接引用):
- "Sonic 3.6 currently leads the Text to Speech Arena with an Elo score of 1272."
- Top 5:1. Sonic 3.6 (1272)、2. Qwen-Audio-3.0-TTS-Plus (1260)、3. Realtime TTS-2 (1245)、4. Simba 3.2 (1237)、5. Luna TTS (1230)。
- "Kokoro 82M v1.0 is the most affordable at $0.65 per 1M characters with an Elo score of 1061. Other affordable options include StyleTTS 2 at $2.82 per 1M characters."
- "Breeze TTS 2 is the highest-ranked open weights model on the Text to Speech Leaderboard with an Elo score of 1204. There are 16 open weights models out of 90 total."
- "The top 5 open weights Text to Speech models are: 1. Breeze TTS 2 (Elo 1204), 2. Fish Audio S2 Pro (Elo 1121), 3. Step Audio EditX (Mar 2026) (Elo 1094), 4. Voxtral TTS (Elo 1075), and 5. Magpie-Multilingual 357M (Feb 2026) (Elo 1063)."
数据来自页面内嵌 payload 中每个模型的默认视图行(同一模型在页面上有多个筛选切片的 Elo,我取默认视图那一条),含 95% 置信区间与投票出现次数:
| Elo | CI95 | appearances | 模型 | 权重来源 URL(页面内 openWeightsUrl) |
|---|---|---|---|---|
| 1204.03 | 1188–1220 | 1,374 | Breeze TTS 2 (BreezeBlue, 2026-08-23) | https://huggingface.co/BreezeBlue/Breeze-TTS-2 |
| 1121.03 | 1108–1134 | 2,205 | Fish Audio S2 Pro (Fish Audio, 2026-03-10, 4B) | https://huggingface.co/fishaudio/s2-pro |
| 1093.76 | — | — | Step Audio EditX (Mar 2026) (StepFun) | https://huggingface.co/stepfun-ai/Step-Audio-EditX |
| 1075.06 | 1062–1088 | 2,028 | Voxtral TTS (Mistral, 2026-03-26) | https://huggingface.co/mistralai/Voxtral-4B-TTS-2603 |
| 1062.54 | — | — | Magpie-Multilingual 357M (Feb 2026) (NVIDIA) | https://huggingface.co/nvidia/magpie_tts_multilingual_357m |
| 1060.98 | 1050–1072 | 5,224 | Kokoro 82M v1.0 (Kokoro, 2025-01-27) | https://huggingface.co/hexgrad/Kokoro-82M |
| 1040.56 | 1021–1061 | 1,681 | OpenAudio S1 Mini (Fish Audio, 2025-06-03) | https://huggingface.co/fishaudio/s1-mini |
| 1039.88 | — | — | Maya1 (Maya Research) | https://huggingface.co/maya-research/maya1 |
| 1032.15 | 1018–1046 | 1,957 | Higgs Audio V3 TTS (Boson AI, 2026-06-04, 4B) | https://huggingface.co/bosonai/higgs-audio-v3-tts-4b |
| 1020.54 | 1009–1033 | 4,595 | Chatterbox (Resemble AI, 2025-05-28) | https://github.com/resemble-ai/chatterbox |
| 1000.00 | 1000–1000 | 4,792 | Zonos-v0.1 (Zyphra, 2025-02-10) | https://github.com/Zyphra/Zonos |
| 950.85 | — | — | VibeVoice 1.5B (Microsoft) | https://huggingface.co/microsoft/VibeVoice-1.5B |
| 946.34 | 932–960 | 2,959 | OpenVoice v2 (2024-04-01) | https://huggingface.co/myshell-ai/OpenVoiceV2 |
| 913.81 | 899–929 | 2,514 | XTTS v2 (Coqui, 2023-11-08) | https://huggingface.co/coqui/XTTS-v2 |
| 891.24 | 876–906 | 2,424 | StyleTTS 2 (2023-06-13) | https://github.com/yl4579/StyleTTS2 |
| 842.56 | — | — | MetaVoice v1 | https://github.com/metavoiceio/metavoice-src |
| rank | Elo | CI95 | appearances | 模型 | 厂商 |
|---|---|---|---|---|---|
| 0 | 1271.98 | 1255–1289 | 1,740 | Sonic 3.6 | Cartesia |
| 1 | 1260.38 | 1243–1277 | 1,438 | Qwen-Audio-3.0-TTS-Plus | Alibaba(页面 url 为 /text-to-speech/model-families/cosyvoice-tts,即归入 CosyVoice 家族;openWeights: false) |
| 2 | 1244.94 | 1227–1263 | 1,229 | Realtime TTS-2 | Inworld |
| 3 | 1236.88 | 1223–1251 | 2,381 | Simba 3.2 | Speechify |
| 4 | 1230.20 | 1216–1244 | 2,502 | Luna TTS | VUI Labs |
| 7 | 1198.97 | 1182–1216 | 1,262 | StepAudio 2.5 TTS (Aug 2026) | StepFun |
| 8 | 1198.86 | 1187–1211 | 3,415 | Gemini 3.1 Flash TTS | |
| 18 | 1137.90 | 1125–1151 | 2,147 | Fish Audio S2.1 Pro | Fish Audio |
| 30 | 1096.35 | 1083–1109 | 2,295 | Chatterbox HD | Resemble AI |
| 76 | 940.30 | — | — | Qwen3 TTS Flash | Alibaba(closed) |
| 79 | 925.98 | 913–939 | 2,854 | Qwen3 TTS | Alibaba(closed) |
榜上没有的(关键空缺,未查到):IndexTTS / IndexTTS-2 / IndexTTS-2.5、CosyVoice 2/3 的开源权重版、GPT-SoVITS、ChatTTS、F5-TTS、MegaTTS3、Spark-TTS、GLM-TTS、FireRedTTS-2、OmniVoice、MOSS-TTS、VoxCPM2、Fish-Speech 的 v1.x 开源版。
注:AA 页面的
Controlled Voice Arena(同样 8 个克隆音色)也内嵌在同一页 payload 中,但我抓到的 payload 里同一模型存在 ~12 个不同的筛选切片 Elo(按 类别×口音 切分),无法确定哪一条是"总榜",因此本报告不引用该板的数字,以免误标。若要引用,必须补抓该板页面并确认筛选状态。
来源:[第三方] 独立研究者 kadirnar 的 VoiceHub Arena,不隶属任何 TTS 厂商。
https://hf-mirror.com/datasets/kadirnar/voicehub-arena-seed-tts-eval(lastModified 2026-09-19;README + leaderboard.csv / leaderboard.json 已下载)https://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval(lastModified 2026-09-19)协议(README 原文,条件非常明确):
en/meta.lst,publisher revision 752f4297f090c46bb1a55a1f7439e5944ddefe8d("The publisher describes the English test material as originating from Common Voice")Systran/faster-whisper-large-v3,rev edaa852e,CUDA FP16,beam 5,temperature 0,无 VAD,无 previous-text conditioning;normalizer whisper-normalizer==0.1.12 / whisper_english完整 33 模型榜(leaderboard.csv,按 WER 升序;我把小数换算成 %)
| # | 模型 | checkpoint | scored | WER%↓ | CER%↓ | exact_match% | RTF | 峰值显存 MiB |
|---|---|---|---|---|---|---|---|---|
| 1 | Kokoro | hexgrad/Kokoro-82M | 1088 | 0.9629 | 0.2440 | 91.36 | 0.019 | 478 |
| 2 | Supertonic | Supertone/supertonic-3 | 1088 | 0.9797 | 0.2663 | 90.99 | 0.064 | 419 |
| 3 | OmniVoice | k2-fsa/OmniVoice | 1088 | 0.9880 | 0.2648 | 91.54 | 0.329 | 3,229 |
| 4 | SpeechT5 | microsoft/speecht5_tts | 1088 | 1.0131 | 0.2946 | 90.63 | 0.162 | 726 |
| 5 | F5-TTS | F5TTS_v1_Base | 1088 | 1.0550 | 0.3199 | 90.90 | 0.782 | 1,516 |
| 6 | StyleTTS2 | styletts2 epochs_2nd_00020 | 1088 | 1.0885 | 0.2991 | 90.44 | 0.042 | 1,068 |
| 7 | EchoTTS | jordand/echo-tts-base | 1088 | 1.1220 | 0.3482 | 90.35 | 0.572 | 9,363 |
| 8 | FishTTS (S2 Pro) | fishaudio/s2-pro | 1088 | 1.1471 | 0.3169 | 89.71 | 3.743 | 14,653 |
| 9 | XTTS | artifacts/xtts2 | 1088 | 1.1722 | 0.3928 | 89.80 | 0.377 | 2,045 |
| 10 | Zonos2 | Zyphra/ZONOS2 | 1088 | 1.1806 | 0.4211 | 90.35 | 6.228 | 15,766 |
| 11 | Chatterbox | ResembleAI/chatterbox | 1088 | 1.2308 | 0.3526 | 88.51 | 0.822 | 3,244 |
| 12 | MOSS-TTS | OpenMOSS-Team/MOSS-TTS-v1.5 | 1088 | 1.2308 | 0.4315 | 90.81 | 1.390 | 25,009 |
| 13 | OuteTTS | Llama-OuteTTS-1.0-1B | 1088 | 1.2392 | 0.4374 | 89.34 | 3.165 | 4,374 |
| 14 | VibeVoice | microsoft/VibeVoice-Realtime-0.5B | 1088 | 1.3146 | 0.4151 | 88.24 | 0.621 | 2,034 |
| 15 | Qwen3-TTS | Qwen3-TTS-12Hz-1.7B-CustomVoice | 1088 | 1.3816 | 0.4389 | 88.33 | 1.770 | 4,351 |
| 16 | NeuTTS | neuphonic/neutts-2e | 1088 | 1.4737 | 0.6204 | 88.88 | 1.704 | 3,476 |
| 17 | CosyVoice 3(修正后) | artifacts/cosyvoice3(=Fun-CosyVoice3-0.5B-2512 base llm.pt) | 1088 | 1.7416 | 0.6234 | 85.48 | 0.703 | 3,430 |
| 18 | HiggsTTS | bosonai/higgs-tts-2-3b-base | 1088 | 1.7416 | 0.6710 | 85.94 | 1.172 | 11,947 |
| 19 | InflectTTS | owensong/Inflect-Micro-v2 | 1088 | 1.7584 | 0.5892 | 85.66 | 0.014 | 174 |
| 20 | OrpheusTTS | canopylabs/orpheus-3b-0.1-ft | 1088 | 2.0012 | 0.7692 | 83.73 | 2.583 | 6,410 |
| 21 | Zonos | Zyphra/Zonos-v0.1-transformer | 1088 | 2.3026 | 0.9641 | 80.24 | 1.768 | 4,145 |
| 22 | GPT-SoVITS | lj1995/GPT-SoVITS | 1088 | 2.6375 | 0.8704 | 82.26 | 0.322 | 764 |
| 23 | MeloTTS | EN | 1088 | 2.6375 | 1.0653 | 77.67 | 0.020 | 507 |
| 24 | ParlerTTS | parler-tts-mini-v1 | 1088 | 2.6543 | 1.4462 | 80.79 | 1.877 | 4,507 |
| 25 | CSM | sesame/csm-1b | 1088 | 2.6710 | 1.4878 | 80.88 | 2.253 | 4,331 |
| 26 | VoxCPM | openbmb/VoxCPM2 | 1088 | 3.3492 | 2.5992 | 83.64 | 0.874 | 5,568 |
| 27 | OpenVoice | myshell-ai/OpenVoiceV2 | 1088 | 3.6842 | 1.5905 | 72.06 | 0.042 | 753 |
| 28 | Vits | facebook/mms-tts-eng | 1088 | 6.2296 | 2.6736 | 60.85 | 0.036 | 162 |
| 29 | Bark | suno/bark-small | 1088 | 7.7367 | 4.9649 | 64.61 | 1.351 | 1,867 |
| 30 | ConversationTTS | artifacts/conversationtts | 1088 | 8.7164 | 8.2470 | 76.56 | 2.672 | 4,321 |
| 31 | Vui | vui-abraham-100m.pt | 1088 | 12.1494 | 8.3973 | 49.63 | 0.412 | 3,021 |
| 32 | Dia | nari-labs/Dia-1.6B-0626 | 1088 | 67.3533 | 59.2826 | 45.22 | 3.189 | 9,101 |
| 33 | Llasa | HKUSTAudio/Llasa-1B-Multilingual | 1088 | 73.9931 | 51.4343 | 3.13 | 0.867 | 5,498 |
质量审计(README 原文,重要):
"CosyVoice is Fun-CosyVoice3-0.5B-2512, base llm.pt. The archived 13.82% WER is affected by a confirmed HiFT implementation defect and is excluded from ranking. The corrected full 1,088-text evaluation is verified: WER 1.7416%, CER 0.6234%." "Llasa and Dia remain under quality review after independent LM/codec checks." "Dia and Llasa have high error rates with unresolved root causes."
同一 33 模型榜里没有的:ChatTTS、MegaTTS3、IndexTTS、Spark-TTS、GLM-TTS、FireRedTTS-2(→ 这几个在该英文基准上 未查到)。
https://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-full-20260916/full/<method>/result.jsongeneration_provenance 读出):NVIDIA RTX 6000 Ada Generation (49,140 MiB),max_gpu_jobs=4,torch 2.8.0+cu128,transformers 4.51.3,soundfile 0.14.0metric_provenance 里 ASR = faster-whisper-large-v3 (cuda/fp16/beam5);另有 dnsmos、utmos22、wavlm_sim(WavLM-large ECAPA speaker SIM)我实际抓取并聚合的 4 个 full-run(1088 条英文 Seed-TTS-Eval,全部 evaluated=1088, generation_failures=0):
| 实验 | repo | WER | CER | exact_match | DNSMOS-OVRL | UTMOS22 | WavLM SIM | RTF | 峰值显存 |
|---|---|---|---|---|---|---|---|---|---|
omnivoice--voice_clone | k2-fsa/OmniVoice | 1.0718% | 0.2782% | 90.81% | 3.1933 | 3.9132 | 0.7385 (n=1088, CI 0.7341–0.7428) | 0.223 | 2,207 MiB |
f5tts--voice_clone | SWivid/F5-TTS | 1.2643% | 0.3853% | 88.69% | 3.0783 | 3.7026 | 0.6691 (n=1088, CI 0.6643–0.6744) | 0.606 | 781 MiB |
kokoro--preset_voice | hexgrad/Kokoro-82M | 0.9462% | 0.2410% | 91.45% | 3.4267 | 4.5037 | N/A(无参考音频) | 0.042 | 480 MiB |
vibevoice--preset_voice_stream | microsoft/VibeVoice-Realtime-0.5B | 1.3481% | 0.4538% | 87.78% | 3.3381 | 4.4490 | N/A(无参考音频) | 0.629 | 2,912 MiB |
⚠️ 注意:同一模型在 3.1 的长表与 3.2 的 campaign 里 WER 略有差异(如 Kokoro 0.9629% vs 0.9462%,OmniVoice 0.9880% vs 1.0718%),因为两个 campaign 是不同的运行批次/方法契约;引用时须带 campaign 名。 另外 README 特别警告:
native-ada-20260916里的 8 条文本 pilot 分数不得当作 full 分数(我下载的native-comparison.csv18 行全部是 pilot,故不引用其数值)。
https://ar5iv.labs.arxiv.org/html/2605.28618(arXiv 2605.28618)[第三方](学术第三方,非模型方):"SwanBench-Speech,1,101 samples spanning 17 common speech scenarios","seven metrics",覆盖 acoustics / semantics / expressiveness 三轴,含 dialog generation,代码/演示 https://swanaigc.github.io/#bench(a) 博客园 sensorsen《语音模型 2026 开源 TTS 选型指南:六款主流语音合成模型实测对比》(2026) [第三方](个人博客,非严谨实测) URL:https://www.cnblogs.com/sensorsen/p/21367537
(b) CSDN《和GPT-SoVITS比如何?两款热门中文TTS横向对比》(2026-05-28) [第三方,低可信/AI 生成痕迹明显] URL:https://blog.csdn.net/weixin_42602241/article/details/156922929
(c) 博客园 kacoro《tts哪家强?》(2025-07-10) [第三方](真实使用日志) URL:https://www.cnblogs.com/kacoro/p/18977125
(d) 腾讯云开发者社区《我们把 Kokoro 1.1 量化到 FP16,并用它替换了浏览器里的中文 Piper 配音》(2026-08-08) [第三方] URL:https://cloud.tencent.cn/developer/article/2722884
#Kokoro1.1 #文字转语音 #onnx(e) SiliconFlow《终极指南 — 2026年最佳开源 Text-to-Speech 模型》 [厂商内容营销] URL:https://www.siliconflow.com/zh/articles/best-open-source-text-to-speech-models
https://www.siliconflow.com/zh/articles/best-open-source-models-for-voice-cloning(本轮未抓取正文)(f) 中文 Arena 的公开数据(未被充分利用的线索) [第三方原始数据]
https://hf-mirror.com/datasets/JacobLinCool/zh-tw-tts-arena-votes(lastModified 2026-07-24,license cc-by-4.0) README 原文:"Append-only log of blind A/B preference votes collected by the zh-tw-tts-arena Space. Schema: ts, session, sentence_id, model_a, model_b, condition_a, condition_b, winner(a|b|tie), dwell_s, ua, app_version."(4 个 jsonl 文件) → 这是本轮唯一找到的「中文(繁體)盲测 Arena」原始投票数据;我只抓了 README,未统计 → 中文 Arena 排名:未查到(可后续补充)(a) The AI Bench — "Best local AI voice models in 2026"(VERIFIED SEPTEMBER 2026) [第三方] URL:https://theaibench.ai/use-cases/voice/
(b) Sovereign AI Blog(2026-05-12) [第三方] — 见 1.2 来源 A;另有重要的工程现实结论:
instructions parameter is silently ignored. The ref_audio parameter crashes the engine because the encoder weights stayed gated in Mistral's hosted product. No speed knob exists."(c) Pinggy / Replicate 等其他英文页:https://pinggy.io/blog/best_open_source_self_hosted_text_to_speech_models/(本轮未抓取正文);https://replicate.com/kjjk10/kokoro-82m/readme(已抓,见 1.3)
先给结论(针对指定的 4 个老模型):
- GPT-SoVITS:Seed-TTS test-zh/test-en 的 WER/SIM 未查到;只查到 CV3-eval 上的 ZH-CER 7.34% / EN-WER 12.5%(VoxCPM 论文,见 4.1);另有 VoiceHub 的英文 WER 2.6375%/CER 0.8704%(见 3.1)。
- MegaTTS3:厂商论文给的是 LibriSpeech-PC 的 SIM-O / WER(英文,非 SeedTTS);SeedTTS 口径的 EN-WER 2.79 / EN-SIM 77.1 / ZH-CER 1.52 / ZH-SIM 79.0 来自他人论文的转引表(见 4.1、4.2)。
- Kokoro:SeedTTS 的 WER/CER/SIM-o 未查到(Kokoro 无官方论文);最接近的是 VoiceHub 的英文 Seed-TTS-Eval WER 0.9629%(见 3.1)与 KaniTTS 的英文 MOS/WER/CER(见 4.4)。
- ChatTTS:SeedTTS 的 WER/CER/SIM 完全未查到(本轮在所有抓取的论文表格中都没有出现 ChatTTS)。
[厂商自评](论文作者是 VoxCPM 方)URL:https://arxiv.org/html/2509.24650v1(arXiv 2509.24650,标题 VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning)
Table 3: Performance on Seed-TTS-eval Benchmark(原文表头:Model | Params | Open-Source | EN WER↓ SIM↑ | ZH CER↓ SIM↑ | Hard CER↓ SIM↑) 所有 SIM 为百分数(0-100),与其它论文的 0-1 口径不同,注意换算:
| Model | Params | OS | EN-WER↓ | EN-SIM↑ | ZH-CER↓ | ZH-SIM↑ | Hard-CER↓ | Hard-SIM↑ |
|---|---|---|---|---|---|---|---|---|
| MegaTTS3 (Jiang et al., 2025) | 0.5B | ✗ | 2.79 | 77.1 | 1.52 | 79.0 | – | – |
| DiTAR | 0.6B | ✗ | 1.69 | 73.5 | 1.02 | 75.3 | – | – |
| CosyVoice3 (0.5B) | 0.5B | ✗ | 2.02 | 71.8 | 1.16 | 78.0 | 6.08 | 75.8 |
| CosyVoice3 (1.5B) | 1.5B | ✗ | 2.22 | 72.0 | 1.12 | 78.1 | 5.83 | 75.8 |
| Seed-TTS | – | ✗ | 2.25 | 76.2 | 1.12 | 79.6 | 7.59 | 77.6 |
| MiniMax-Speech | – | ✗ | 1.65 | 69.2 | 0.83 | 78.3 | – | – |
| F5-TTS | 0.3B | ✓ | 2.00 | 67.0 | 1.53 | 76.0 | 8.67 | 71.3 |
| MaskGCT | – | ✓ | 2.62 | 71.7 | 2.27 | 77.4 | – | – |
| CosyVoice | 0.3B | ✓ | 4.29 | 60.9 | 3.63 | 72.3 | 11.75 | 70.9 |
| CosyVoice2 | 0.5B | ✓ | 3.09 | 65.9 | 1.38 | 75.7 | 6.83 | 72.4 |
| SparkTTS | 0.5B | ✓ | 3.14 | 57.3 | 1.54 | 66.0 | – | – |
| FireRedTTS | 0.5B | ✓ | 3.82 | 46.0 | 1.51 | 63.5 | 17.45 | 62.1 |
| FireRedTTS-2 | – | ✓ | 1.95 | 66.5 | 1.14 | 73.6 | – | – |
| Qwen2.5-Omni | 7B | ✓ | 2.72 | 63.2 | 1.70 | 75.2 | 7.97 | 74.7 |
| OpenAudio-s1-mini | 0.5B | ✓ | 1.94 | 55.0 | 1.18 | 68.5 | 23.37 | 64.3 |
| IndexTTS 2 | 1.5B | ✓ | 2.23 | 70.6 | 1.03 | 76.5 | 7.12 | 75.5 |
| VibeVoice | 1.5B | ✓ | 3.04 | 68.9 | 1.16 | 74.4 | – | – |
| HiggsAudio-v2 | 3B | ✓ | 2.44 | 67.7 | 1.50 | 74.0 | 55.07 | 65.6 |
| VoxCPM-Emilia | 0.5B | ✓ | 2.34 | 68.1 | 1.11 | 74.0 | 12.46 | 69.8 |
| VoxCPM | 0.5B | ✓ | 1.85 | 72.9 | 0.93 | 77.2 | 8.87 | 73.0 |
Table 4: Performance on CV3-eval Benchmark(*denotes close-sourced systems)——这是唯一找到的 GPT-Sovits 数字:
| Model | CV3-EVAL ZH-CER↓ | CV3-EVAL EN-WER↓ | CV3-Hard-ZH CER↓ | SIM↑ | DNSMOS↑ | CV3-Hard-EN WER↓ | SIM↑ | DNSMOS↑ |
|---|---|---|---|---|---|---|---|---|
| F5-TTS | 5.47 | 8.90 | – | – | – | – | – | – |
| SparkTTS | 5.15 | 11.0 | – | – | – | – | – | – |
| GPT-Sovits | 7.34 | 12.5 | – | – | – | – | – | – |
| CosyVoice2 | 4.08 | 6.32 | 12.58 | 72.6 | 3.81 | 11.96 | 66.7 | 3.95 |
| OpenAudio-s1-mini | 4.00 | 5.54 | 18.1 | 58.2 | 3.77 | 12.4 | 55.7 | 3.89 |
| IndexTTS2 | 3.58 | 4.45 | 12.8 | 74.6 | 3.65 | 8.78 | 74.5 | 3.80 |
| HiggsAudio-v2 | 9.54 | 7.89 | 41.0 | 60.2 | 3.39 | 10.3 | 61.8 | 3.68 |
| CosyVoice3-0.5B* | 3.89 | 5.24 | 14.15 | 78.6 | 3.75 | 9.04 | 75.9 | 3.92 |
| CosyVoice3-1.5B* | 3.91 | 4.99 | 9.77 | 78.5 | 3.79 | 10.55 | 76.1 | 3.95 |
| VoxCPM-Emilia | 4.47 | 5.23 | 22.2 | 62.6 | 3.47 | 10.00 | 62.6 | 3.68 |
| VoxCPM | 3.40 | 4.04 | 12.9 | 66.1 | 3.59 | 7.89 | 64.3 | 3.74 |
口径警告:CV3-eval 的 CER 数值(4~8%)远高于 Seed-TTS-eval test-zh 的 CER(1~2%),两套测试集的正文长度/难度不同。不能用 CV3-eval 的 7.34% 去和 test-zh 的 1.5% 对比。
[厂商自评]URL:https://arxiv.org/html/2502.18924v1(标题 Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis,即 MegaTTS3 / S-DiT)
主表(Table 5,英文 LibriSpeech-PC 类测试集):
| Model | #Params | Training Data | SIM-O↑ | SIM-R↑ | WER↓ | CMOS↑ | SMOS↑ | RTF↓ |
|---|---|---|---|---|---|---|---|---|
| GT | – | – | 0.68 | – | 1.94% | +0.12 | 3.92 | – |
| VALL-E 2* | 0.4B | LibriHeavy | 0.64 | 0.68 | 2.44% | – | – | – |
| VoiceBox† | 0.4B | Collected (60kh) | 0.64 | 0.67 | 2.03% | −0.20 | 3.81 | 0.340 |
| DiTTo-TTS* | 0.7B | Collected (55kh) | 0.62 | 0.65 | 2.56% | – | – | – |
| NaturalSpeech 3† | 0.5B | LibriLight | 0.67 | 0.76 | 1.81% | −0.10 | 3.95 | 0.296 |
| CosyVoice | 0.4B | Collected (172kh) | 0.62 | – | 2.24% | −0.18 | 3.93 | 1.375 |
| MaskGCT | 1.0B | Emilia (100kh) | 0.69 | – | 2.63% | – | – | – |
| F5-TTS | 0.3B | Emilia (100kh) | 0.66 | – | 1.96% | −0.12 | 3.96 | 0.307 |
| S-DiT(=MegaTTS3) | 0.3B | LibriLight | 0.71 | 0.78 | 1.82% | 0.00 | 3.98 | 0.188 |
| S-DiT-accelerated | 0.3B | LibriLight | 0.70 | 0.78 | 1.86% | −0.03 | 3.96 | 0.124 |
其余表(同一论文):Table 6 SIM-O 0.70 / WER 2.31%;Table 11 Ours 0.71 / 1.82%;Table 12 规模消融(0.5B: SIM-O 0.66 / WER 2.10%;1.5B: 0.72 / 1.98%;7.0B: 0.74 / 1.90%);Table 14 长文本 WER 2.39%(vs CosyVoice 5.52%、VoiceCraft 12.81%)。
(a) RobustSpeechFlow(arXiv 2605.22083)的 Table 7 [第三方学术] URL:https://arxiv.org/html/2605.22083v1(RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching)
表头:
Model | Params | WER↓ | SIM↑MegaTTS3 [23] | 0.5B | 2.79 | 0.77;Seed-TTS DiT [21] | – | 1.73 | 0.79;DiTAR [5] | 0.6B | 1.69 | 0.74;MiniMax-Speech [24] | – | 1.65 | 0.69;F5-TTS [2] | 0.3B | 2.00 | 0.67;CosyVoice3 [25] | 1.5B | 2.22 | 0.72;Spark-TTS [26] | 0.5B | 3.14 | 0.57;OpenAudio S1-Mini [27] | 0.5B | 1.94 | 0.55;IndexTTS2 [28] | 1.5B | 2.23 | 0.71;VibeVoice [7] | 1.5B | 3.04 | 0.69;VoxCPM-Emilia [6] | 0.5B | 2.34 | 0.68;VoxCPM [6] | 0.5B | 1.85 | 0.73
(b) WavTTS(arXiv 2606.03455)Table 10 —— Seed-TTS test-en / test-zh [第三方学术] URL:https://arxiv.org/html/2606.03455v1(WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling) 表头:Model | Params | Data(hrs) | Seed-TTS test-en: WER%↓ SIM-o↑ UTMOS↑ | Seed-TTS test-zh: CER%↓ SIM-o↑ UTMOS↑
| Model | Params | Data | en WER% | en SIM-o | en UTMOS | zh CER% | zh SIM-o | zh UTMOS |
|---|---|---|---|---|---|---|---|---|
| Ground Truth | – | – | 1.79 | 0.73 | 3.53 | 1.25 | 0.75 | 2.78 |
| CosyVoice | 416M | 170K Multi. | 4.29 | 0.61 | – | 3.63 | 0.72 | – |
| CosyVoice 2 | 618M | 167K Multi. | 2.57 | 0.65 | – | 1.45 | 0.75 | – |
| Llasa-1B | 1370M | 250K Multi. | 3.22 | 0.57 | – | 1.89 | 0.67 | – |
| Spark-TTS | 507M | 102K Multi. | 1.98 | 0.58 | – | 1.20 | 0.67 | – |
| MaskGCT | 1048M | 100K Emilia | 2.36 | 0.71 | 3.57 | 2.48 | 0.77 | 2.64 |
| E2-TTS | 333M | 100K Emilia | 2.21 | 0.71 | 3.20 | 1.97 | 0.73 | 2.27 |
| F5-TTS | 336M | 100K Emilia | 1.65 | 0.66 | 3.73 | 1.55 | 0.75 | 2.94 |
| ZipVoice | 123M | 100K Emilia | 1.60 | 0.70 | 3.83 | 1.40 | 0.75 | 3.15 |
| LongCat-AudioDiT | 1420M | 100K Multi. | 1.94 | 0.76 | 3.80 | 1.10 | 0.81 | 3.16 |
| WavTTS | 673M | 100K Emilia | 1.50 | 0.65 | 3.92 | 1.59 | 0.73 | 3.08 |
(无 GPT-SoVITS / MegaTTS3 / Kokoro / ChatTTS)
(c) CaT-TTS(arXiv 2509.22062)Table 11 [第三方学术] URL:https://arxiv.org/html/2509.22062v1(Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling) 表头 Model | test-zh WER%↓ SIM↑ | test-en WER%↓ SIM↑ | test-hard WER%↓ SIM↑:
| Model | zh WER% | zh SIM | en WER% | en SIM | hard WER% | hard SIM |
|---|---|---|---|---|---|---|
| MaskGCT | 2.27 | 0.774 | 2.62 | 0.714 | 10.27 | 0.748 |
| E2 TTS (32 NFE) | 1.97 | 0.730 | 2.19 | 0.710 | – | – |
| F5-TTS (32 NFE) | 1.56 | 0.741 | 1.83 | 0.647 | 8.67 | 0.713 |
| Seed-TTS | 1.12 | 0.796 | 2.25 | 0.762 | 7.59 | 0.776 |
| FireRedTTS | 1.51 | 0.635 | 3.82 | 0.460 | 17.45 | 0.621 |
| CosyVoice | 3.63 | 0.723 | 4.29 | 0.609 | 11.75 | 0.709 |
| CosyVoice 2 | 1.45 | 0.748 | 2.57 | 0.652 | 6.83 | 0.724 |
| CosyVoice 3-0.5B | 1.16 | 0.780 | 2.02 | 0.718 | 6.08 | 0.758 |
| QTTS | 1.66 | 0.648 | 3.17 | 0.652 | 14.45 | 0.641 |
| Spark-TTS | 1.20 | 0.672 | 1.98 | 0.584 | – | – |
| Llasa-1B/3B/8B-250k | 1.89/1.60/1.59 | 0.668/0.675/0.684 | 3.22/3.14/2.97 | 0.572/0.579/0.574 | 12.13/13.37/11.09 | 0.638/0.652/0.660 |
| CaT-TTS | 1.56 | 0.678 | 2.35 | 0.668 | 9.75 | 0.674 |
(无 GPT-SoVITS / MegaTTS3 / Kokoro / ChatTTS)
https://hf-mirror.com/api/models/hexgrad/Kokoro-82M 返回 tags 含 "text-to-speech","en" → 官方仅英文。https://hf-mirror.com/api/models?search=Kokoro,按下载量排序):
hexgrad/Kokoro-82M-v1.1-zh(downloads 22,373)onnx-community/Kokoro-82M-v1.1-zh-ONNX(tags 含 en, zh)aufklarer/Kokoro-82M-CoreML(tags 含 en, zh)contextboxai/Kokoro-Vietnamese、zaakirio/kokoro-ru、Thorsten-Voice/Kokoro(de)、TilLabs/kokoro-tts-kazakh 等https://github.com/Lyrcaxis/KokoroSharp/issues/5(2025-02-11,issue 正文):"I tried this library to generate Chinese speech, but I found that it has a very heavy accent, and it's a bit hard to hear what's being said in the generated content, but using the python version of the kokoro library, the generated Chinese speech files don't have this problem."(即问题出在 C# 端实现,而非模型本身)[第三方]
https://hf-mirror.com/datasets/KaniTTS-research-team/kokoro_banchmark(lastModified 2026-03-18;1 split,3600 行,7.2 MB parquet,含 MOS/NOI/DIS/COL/LOUD/UTMOS/WER/CER)| model(Kokoro 英文音色) | n | MOS | UTMOS | WER | CER |
|---|---|---|---|---|---|
| kokoro_82m_en_am_fenrir | 720 | 4.986 | 3.760 | 0.013 | 0.005 |
| kokoro_82m_en_af_heart | 720 | 4.981 | 3.759 | 0.013 | 0.004 |
| kokoro_82m_en_am_michael | 720 | 4.923 | 3.814 | 0.012 | 0.004 |
| kokoro_82m_en_af_bella | 720 | 4.538 | 3.122 | 0.010 | 0.004 |
| kokoro_82m_en_bf_emma | 720 | 4.453 | 3.501 | 0.012 | 0.004 |
(MOS ≈ 4.45–4.99 偏高,疑似非严格人工 MOS;WER/CER 为比例值,×100 即百分数:1.0–1.3% WER、0.4–0.5% CER。注意这是英文。)
(a) CosyVoice README 官方表(Alibaba FunAudioLLM) [厂商自评] URL:https://raw.githubusercontent.com/FunAudioLLM/CosyVoice/main/README.md 表头:Model | Open-Source | Size | test-zh CER%↓ SS%↑ | test-en WER%↓ SS%↑ | test-hard CER%↓ SS%↑
| Model | OS | Size | test-zh CER | test-zh SS | test-en WER | test-en SS | test-hard CER | test-hard SS |
|---|---|---|---|---|---|---|---|---|
| Human | – | – | 1.26 | 75.5 | 2.14 | 73.4 | – | – |
| Seed-TTS | ❌ | – | 1.12 | 79.6 | 2.25 | 76.2 | 7.59 | 77.6 |
| MiniMax-Speech | ❌ | – | 0.83 | 78.3 | 1.65 | 69.2 | – | – |
| F5-TTS | ✅ | 0.3B | 1.52 | 74.1 | 2.00 | 64.7 | 8.67 | 71.3 |
| Spark TTS | ✅ | 0.5B | 1.2 | 66.0 | 1.98 | 57.3 | – | – |
| CosyVoice2 | ✅ | 0.5B | 1.45 | 75.7 | 2.57 | 65.9 | 6.83 | 72.4 |
| FireRedTTS2 | ✅ | 1.5B | 1.14 | 73.2 | 1.95 | 66.5 | – | – |
| Index-TTS2 | ✅ | 1.5B | 1.03 | 76.5 | 2.23 | 70.6 | 7.12 | 75.5 |
| VibeVoice-1.5B | ✅ | 1.5B | 1.16 | 74.4 | 3.04 | 68.9 | – | – |
| VibeVoice-Realtime | ✅ | 0.5B | – | – | 2.05 | 63.3 | – | – |
| HiggsAudio-v2 | ✅ | 3B | 1.50 | 74.0 | 2.44 | 67.7 | – | – |
| VoxCPM | ✅ | 0.5B | 0.93 | 77.2 | 1.85 | 72.9 | 8.87 | 73.0 |
| GLM-TTS | ✅ | 1.5B | 1.03 | 76.1 | – | – | – | – |
| GLM-TTS RL | ✅ | 1.5B | 0.89 | 76.4 | – | – | – | – |
| Fun-CosyVoice3-0.5B-2512 | ✅ | 0.5B | 1.21 | 78.0 | 2.24 | 71.8 | 6.71 | 75.8 |
| Fun-CosyVoice3-0.5B-2512_RL | ✅ | 0.5B | 0.81 | 77.4 | 1.68 | 69.5 | 5.44 | 75.0 |
(无 GPT-SoVITS / MegaTTS3 / Kokoro / ChatTTS)
(b) IndexTTS 官方 README(bilibili) [厂商自评] URL:https://raw.githubusercontent.com/index-tts/index-tts/main/README.md(同表亦可从 https://raw.githubusercontent.com/T8mars/indextts25-desktop-t8/main/README.md 得到同内容并含 Table 2)
"Table 1: Zero-shot TTS on CV3-Eval (Arabic uses an in-house test set). †Cited from the original paper."
| Model | Params | zh WER% | zh SS% | en WER% | en SS% | es WER% | es SS% | ja WER% | ja SS% | ar WER% | ar SS% | Avg WER% | Avg SS% |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| VoxCPM2 | 2B | 3.88 | 74.99 | 5.13 | 71.57 | 5.49 | 74.67 | 6.69 | 72.90 | 14.94 | 65.99 | 7.22 | 72.02 |
| OmniVoice | 0.8B | 3.41 | 72.99 | 3.62 | 70.13 | 3.52 | 74.14 | 5.38 | 70.49 | 17.88 | 64.22 | 6.76 | 70.39 |
| Moss-TTS 1.5 | 8B | 4.02 | 72.68 | 4.45 | 67.46 | 3.83 | 71.75 | 10.97 | 68.71 | 23.71 | 62.21 | 9.40 | 68.56 |
| CosyVoice3-0.5B | 0.5B | 3.84 | 80.01 | 4.88 | 74.16 | 4.04 | 78.85 | – | 76.36 | – | – | – | – |
| CosyVoice3-1.5B | 1.5B | 3.91† | – | 4.99† | – | 4.47† | – | 7.57† | – | – | – | – | – |
| FireRedTTS-2 | 1.5B | 8.22 | 68.10 | 14.92 | 56.93 | – | – | – | – | – | – | – | – |
| Fish Audio S2 Pro | 4B | 3.62 | 67.79 | 3.83 | 61.66 | 2.93 | 67.44 | 5.15 | 66.15 | 14.15 | 59.43 | 5.94 | 64.49 |
| Qwen3-TTS | 1.7B | 3.27 | 73.02 | 5.06 | 67.17 | 2.87 | 73.17 | 5.89 | 70.18 | – | – | – | – |
| IndexTTS2.5 | 0.8B | 4.36 | 77.10 | 5.12 | 68.06 | 3.75 | 76.39 | 5.66 | 74.62 | 14.88 | 69.74 | 6.75 | 73.18 |
| IndexTTS2.5-RL | 0.8B | 3.93 | 77.92 | 3.89 | 67.79 | 3.33 | 76.68 | 5.30 | 75.41 | 13.58 | 70.36 | 6.00 | 73.63 |
Table 2(跨语言,中文 prompt → 目标语言)关键行:OmniVoice zh→en WER 3.74 / SS 64.91;zh→es 5.84 / 62.08;zh→ja 9.09 / 69.06;zh→ar 19.80 / 65.27;平均 9.62 / 65.33。IndexTTS2.5-RL 平均 6.17 / 70.20。 (注:这套表是 IndexTTS 团队做的,把 OmniVoice / VoxCPM2 / CosyVoice3 / Fish S2 Pro / Qwen3-TTS 都排了进去 → 属于厂商对比,不是第三方中立评测;其中带 † 的来自原论文,其余应为该团队自测。)
(c) VoxCPM 官方 README(含 zh 版) [厂商自评] — URL:https://raw.githubusercontent.com/OpenBMB/VoxCPM/main/README.md 提供 Seed-TTS-eval 表(VoxCPM2 2B:zh 1.84 / SIM 75.3;en 0.97 / SIM 79.5;hard 8.13 / 75.3;FishAudio S2 4B:zh 0.99、en 0.54;LongCat-Audio-DiT 3.5B:zh 1.50 / SIM 78.6、en 1.09 / 81.8;Qwen3-TTS 1.7B:zh 1.23、en 1.22;MOSS-TTS:zh 1.85 / SIM 73.4、en 1.20 / 78.8) 及 CV3-eval 多语表(VoxCPM2: zh 3.65 / en 5.00 / hard-zh 8.55 / hard-en 8.48;Fish Audio S2: 2.65 / 2.43 / 9.10 / 4.40) 及 MiniMax-Multilingual-Test(Chinese: FishAudio S2 0.730、Qwen3-TTS 0.928、VoxCPM2 1.136、MiniMax 2.252、ElevenLabs 16.026)
(d) OmniVoice README(k2-fsa) [厂商自评] — URL:https://raw.githubusercontent.com/k2-fsa/OmniVoice/main/README.md
"Benchmark (seed-tts zh testset, 2020 samples / 3.3h audio, voice cloning, single H100, fp16, num_step=32; Average RTF as reported by
omnivoice-infer-batch, outputs ASR-verified lossless)"
| batch | baseline RTF | FlashInfer RTF | speedup |
|---|---|---|---|
| 1 | 0.0899 | 0.0430 | 2.1× |
| 1 + CUDA graph | – | 0.0367 | 2.4× |
| 2 | 0.0480 | 0.0245 | 2.0× |
| 4 | 0.0331 | 0.0152 | 2.2× |
| 8 | 0.0298 | 0.0115 | 2.6× |
arXiv:2604.00688(OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models)(e) Fish Speech README(fishaudio) [厂商自评] — URL:https://raw.githubusercontent.com/fishaudio/fish-speech/main/README.md
https://github.com/vllm-project/vllm-omni/blob/main/recipes/fishaudio/Fish-Speech-S2-Pro.md)| 框架 | 是否需 flash-attn | sm_120 现成 wheel? | 已知报错(原文) | 可行 workaround | 证据 URL |
|---|---|---|---|---|---|
| GPT-SoVITS | 否(核心路径不用) | 是(用 cu128 官方 torch wheel 即可) | NVIDIA GeForce RTX 5070 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation. The current PyTorch install supports CUDA capabilities sm_37 sm_50 sm_60 sm_61 sm_70 sm_75 sm_80 sm_86 sm_90 compute_37;RuntimeError: "no kernel image is available" on new NVIDIA GPUs (e.g., RTX 5090 / sm_120);5060 Ti 上 s2_train.py 的 mp.spawn 报错 | 换 cu128 的 torch(≥2.7);整合包自带旧 torch 是根因。另有 PR #2774「Fix s1_train DDP crash on Windows single-GPU (sm_120 / Blackwell)」 | 见 5.2 |
| CosyVoice / Fun-CosyVoice3 | 否(官方路径用 SDPA;有 triton-TRTLLM 可选) | 是 | Error A: 'torch.distributed_c10d.ProcessGroup' object has no attribute 'options'(torch 2.4+ 移除);Error B(按 requirements 降到 torch 2.3.1 时):NVIDIA GeForce RTX 5060 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation (supports up to sm_90). Note: RTX 50-series cards require modern CUDA toolkits and PyTorch versions not covered by the current requirements.;Error C: ERROR: Cannot install hyperpyyaml==1.2.2 and ruamel.yaml<=0.17.21 because these package versions have conflicting dependencies | 升级到 torch 2.7/2.8+cu128,不要按老 requirements 降级;ruamel.yaml 冲突需手工解 | https://github.com/QwenAudio/CosyVoice/issues/1815 |
| ChatTTS(jianchang512/ChatTTS-ui 整合包) | 否 | 是(但整合包内置旧 torch) | NVIDIA GeForce RTX 5070 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation…(整合包换 50 系卡后运行 app.exe 即报) | 换掉整合包内 torch 为 cu128 版 | https://github.com/jianchang512/ChatTTS-ui/issues/290 |
| Fish-Speech / Fish Audio S2 Pro | 是(上游配方在 sm_120 会坏) | 需用预置容器 | "the standard way to add S2-Pro's DAC codec to it breaks on new RTX 5090 / RTX PRO 6000 Blackwell cards (FlashAttention-3 ships kernels only for Hopper sm90a / Blackwell-Ultra sm120a, and SGLang is blocked on sm_120)" | 第三方项目 Genesis1231/fish-s2-rtx:容器 pinned 到 vLLM-Omni 0.22 的 torch 2.11 + cu130 + sm_120 kernels;另注"Prefix caching is intentionally off — it intermittently produces empty audio in vLLM-Omni 0.22" | https://github.com/Genesis1231/fish-s2-rtx |
| index-tts / IndexTTS2 | 未查到确切结论 | – | 有 issue「能用RTX50系显卡跑吗?」(#242,本轮抓取超时,正文未取得 → 未查到);另有 torch.AcceleratorError: CUDA error: operation not permitted when stream is capturing (#677) | 未查到 | issue 列表来自 GitHub 搜索 API 结果:https://github.com/index-tts/index-tts/issues/242(正文未获取) |
| flash-attn(通用) | — | 有可用的免编译组合 | "many people struggle with building (or installing) flash attention" | 实测可行:Python=3.10 + CUDA=12.8 + torch 2.7.1(pip install torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1 --index-url https://download.pytorch.org/whl/cu128)然后 pip install flash_attn==2.8.0.post2 torch==2.7.1 --no-build-isolation | https://github.com/Dao-AILab/flash-attention/issues/2016 |
| F5-TTS | – | – | 未查到(sneekes.app 的 "F5-TTS Installation Guide for RTX 5070 on WSL2" 两次抓取均超时,正文未取得) | 未查到 | 线索 URL(未成功抓取,不得引用其内容):https://sneekes.app/posts/f5-tts-installation-guide-for-rtx-5070-on-wsl2/ |
| MegaTTS3 / MOSS-TTS | – | – | 未查到 | 未查到 | — |
https://github.com/RVC-Boss/GPT-SoVITS/issues/2205「50系列显卡不支持」(opened 2025-03-19,Closed) 训练时报:NVIDIA GeForce RTX 5070 Ti with CUDA capability sm_120 is not compatible with the current PyTorch installation. The current PyTorch install supports CUDA capabilities sm_37 sm_50 sm_60 sm_61 sm_70 sm_75 sm_80 sm_86 sm_90 compute_37. If you want to use the NVIDIA GeForce RTX 5070 Ti GPU with PyTorch, please check the instructions at https://pytorch.org/get-started/locally/https://github.com/RVC-Boss/GPT-SoVITS/issues/2514「[Bug Report] RuntimeError: "no kernel image is available" on new NVIDIA GPUs (e.g., RTX 5090 / sm_120)」(2025-07-09,状态 In follow-up) 报告人自述:"it seems the current PyTorch binaries do not yet support the architecture of the latest GPUs… I've attempted to fix this by upgrading PyTorch to the latest available versions, including the nightly builds, but it appears that even the newest pre-release binaries do not yet include support for sm_120. The command I used was: pip install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/cu121" → ⚠️ 注意:该用户装的是 cu121 nightly,而 sm_120 需要 cu128 及以上,所以他的尝试方式本身就是错的(这条记录很好用来说明"为什么很多人以为不支持")。https://github.com/RVC-Boss/GPT-SoVITS/issues/2192「NVIDIA RTX 5070 with CUDA 12.8 (sm_120), Error Message」(Closed) 在 GPT_SoVITS/inference_webui.py 阶段出现 transformers/pytree 的 FutureWarning 之后报错(正文含完整 traceback)。https://github.com/RVC-Boss/GPT-SoVITS/issues/2393「搭载 5060 ti显卡,sovits模型训练错误」/ #2394「搭载5060 ti显卡上 训练sovits模型出现的报错。gpt模型是可以训练的」(均 Closed) traceback 指向 GPT_SOVITS/s2_train.py 第 601 行 main() → 第 57 行 mp.spawn(...) → torch/multiprocessing/spawn.py。 → 关键信息:5060 Ti 上「GPT 模型可以训练,SoVITS 模型训练报错」(#2394 标题直述)。https://github.com/RVC-Boss/GPT-SoVITS/pull/2774「Fix s1_train DDP crash on Windows single-GPU (sm_120 / Blackwell)」(open)——说明 DDP 在 sm_120 上仍有坑。--enforce-eager Is Not Enough"(2026-04-25,容器起不来)AttributeError(Python init-order bug)that masqueraded as a Blackwell GPU hang"(2026-05-03)microsoft/VibeVoice-Realtime-0.5B 有"verified setup thread for DGX Spark with CUDA 13 and aarch64",是候选里唯一有公开 Spark 部署记录的引擎https://sovgrid.org/blog/strategy-tts-pivot-voxtral-ceiling/https://github.com/ggml-org/ggml/issues/1466「ggml-cuda: flash-attn MMA picker has no Blackwell (sm_120) entry — silently uses Ampere config」(本报告仅通过搜索发现该 URL,未抓取正文 → 内容未核实)JacobLinCool/zh-tw-tts-arena-votes 的原始 jsonl(未统计)。https://hf-mirror.com/spaces/TTS-AGI/TTS-Arena-V2(镜像拒绝页)https://hf-mirror.com/spaces/TTS-AGI/TTS-Arena-V2/raw/main/README.md(176 B)https://huggingface.co/spaces/TTS-AGI/TTS-Arena-V2(官方域名,超时不可达)https://ttsarena.org、https://ttsarena.org/leaderboard、https://ttsarena.org/api/leaderboard(均为静态 TTS-AGI 项目页)https://docs.ttsarena.orghttps://github.com/TTS-AGI/TTS-Arenahttps://raw.githubusercontent.com/TTS-AGI/TTS-Arena/main/README.mdhttps://artificialanalysis.ai/text-to-speechhttps://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice/open-weightshttps://artificialanalysis.ai/text-to-speech/leaderboard/provider-voicehttps://artificialanalysis.ai/api/text-to-speech/leaderboard(404)https://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-05-23/voice-leaderboards.jsonlhttps://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-06-02/voice-leaderboards.jsonlhttps://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-06-18/voice-leaderboards.jsonlhttps://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-08-01/voice-leaderboards.jsonlhttps://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-09-01/voice-leaderboards.jsonlhttps://hf-mirror.com/datasets/tensorfeed/ai-ecosystem-daily/raw/main/2026-09-20/voice-leaderboards.jsonlhttps://hf-mirror.com/datasets/Pendrokar/TTS_Arena/raw/main/README.mdhttps://hf-mirror.com/datasets/Pendrokar/TTS_Arena/raw/main/tts_arena_vote_summary_all.tsvhttps://hf-mirror.com/api/datasets/Pendrokar/TTS_Arenahttps://replicate.com/kjjk10/kokoro-82m/readmehttps://hf-mirror.com/datasets/kadirnar/voicehub-arena-seed-tts-eval/raw/main/README.mdhttps://hf-mirror.com/datasets/kadirnar/voicehub-arena-seed-tts-eval/raw/main/leaderboard.csvhttps://hf-mirror.com/api/datasets/kadirnar/voicehub-arena-seed-tts-evalhttps://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/README.mdhttps://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-20260916/native-comparison.csvhttps://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-full-20260916/full/omnivoice--voice_clone/result.jsonhttps://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-full-20260916/full/f5tts--voice_clone/result.jsonhttps://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-full-20260916/full/kokoro--preset_voice/result.jsonhttps://hf-mirror.com/datasets/VoiceHub/voicehub-arena-seed-tts-eval/raw/main/experiments/native-ada-full-20260916/full/vibevoice--preset_voice_stream/result.jsonhttps://hf-mirror.com/api/datasets/VoiceHub/voicehub-arena-seed-tts-evalhttps://hf-mirror.com/datasets/KaniTTS-research-team/kokoro_banchmark/raw/main/README.mdhttps://hf-mirror.com/api/datasets/KaniTTS-research-team/kokoro_banchmarkhttps://hf-mirror.com/datasets/KaniTTS-research-team/kokoro_banchmark/resolve/main/data/train-00000-of-00001.parquethttps://hf-mirror.com/datasets/JacobLinCool/zh-tw-tts-arena-votes/raw/main/README.mdhttps://hf-mirror.com/api/datasets?search=tts+arena&limit=30、https://hf-mirror.com/api/datasets?search=TTS-AGI&limit=30https://hf-mirror.com/api/models/hexgrad/Kokoro-82Mhttps://hf-mirror.com/api/models?search=Kokoro&limit=40&sort=downloads&direction=-1https://hf-mirror.com/api/datasets/VoiceHub/voicehub-arena-seed-tts-eval(文件清单)https://hf-mirror.com/api/datasets/kadirnar/voicehub-arena-seed-tts-eval(文件清单)https://arxiv.org/html/2509.24650v1(VoxCPM;含 Table 3 Seed-TTS-eval、Table 4 CV3-eval 及 GPT-Sovits)https://arxiv.org/html/2502.18924v1(MegaTTS3 / S-DiT 官方论文,英文 LibriSpeech-PC SIM-O/WER)https://arxiv.org/html/2509.22062v1(CaT-TTS,SeedTTS test-zh/en/hard)https://arxiv.org/html/2605.22083v1(RobustSpeechFlow,含 MegaTTS3 转引)https://arxiv.org/html/2606.03455v1(WavTTS,Seed-TTS test-en/zh SIM-o)https://ar5iv.labs.arxiv.org/html/2605.28618(SwanBench-Speech 长语音基准)https://raw.githubusercontent.com/FunAudioLLM/CosyVoice/main/README.mdhttps://raw.githubusercontent.com/k2-fsa/OmniVoice/main/README.mdhttps://raw.githubusercontent.com/fishaudio/fish-speech/main/README.mdhttps://raw.githubusercontent.com/index-tts/index-tts/main/README.mdhttps://raw.githubusercontent.com/T8mars/indextts25-desktop-t8/main/README.md(IndexTTS 官方 Table 1/Table 2 全文)https://raw.githubusercontent.com/OpenBMB/VoxCPM/main/README.mdhttps://github.com/T8mars/indextts25-desktop-t8https://github.com/FunAudioLLM/CV3-Eval(页面已抓,本报告未引用其内容)https://github.com/RVC-Boss/GPT-SoVITS/issues/2514https://github.com/RVC-Boss/GPT-SoVITS/issues/2192https://github.com/RVC-Boss/GPT-SoVITS/issues/2205https://github.com/RVC-Boss/GPT-SoVITS/issues/2393https://github.com/RVC-Boss/GPT-SoVITS/pull/2774(经 GitHub 搜索 API 结果发现;正文未单独抓取)https://github.com/RVC-Boss/GPT-SoVITS/issues/2626(经搜索发现;正文未抓取)https://github.com/QwenAudio/CosyVoice/issues/1815https://github.com/jianchang512/ChatTTS-ui/issues/290https://github.com/Dao-AILab/flash-attention/issues/2016https://github.com/Genesis1231/fish-s2-rtxhttps://api.github.com/search/issues?q=repo:RVC-Boss/GPT-SoVITS+sm_120(搜索 API,仅用于发现;后触发速率限制)https://api.github.com/search/issues?q=repo:RVC-Boss/GPT-SoVITS+5060https://api.github.com/search/issues?q=repo:index-tts/index-tts+5060https://huggingface.co/2Noise/ChatTTS/discussions/38(仅搜索结果中出现,未抓取)https://sovgrid.org/blog/strategy-tts-pivot-voxtral-ceiling/https://theaibench.ai/use-cases/voice/https://www.cnblogs.com/sensorsen/p/21367537https://blog.csdn.net/weixin_42602241/article/details/156922929https://www.cnblogs.com/kacoro/p/18977125https://www.siliconflow.com/zh/articles/best-open-source-text-to-speech-modelshttps://cloud.tencent.cn/developer/article/2722884https://github.com/Lyrcaxis/KokoroSharp/issues/5https://pinggy.io/blog/best_open_source_self_hosted_text_to_speech_models/(未抓取正文)https://sneekes.app/posts/f5-tts-installation-guide-for-rtx-5070-on-wsl2/(两次抓取均超时,内容未取得)https://huggingface.co/spaces/Pendrokar/TTS-Spaces-Arena(README 中引用的原 Space 地址;本网络不可达)web_search 工具(用于发现候选 URL,共 6 轮):包括 "TTS Arena leaderboard Elo 2026"、"2026 开源 TTS 对比 实测"、"GPT-SoVITS ChatTTS MegaTTS3 test-zh CER SIM-o comparison table paper"、"CosyVoice RTX 5090 sm_120 flash-attn issue github" 等查询。http://export.arxiv.org/api/query(arXiv API,用于检索候选论文)/tmp/dh/f.sh(curl 抓取)、/tmp/dh/txt.sh(HTML→文本)、自写 tab.py(arXiv 表格抽取)、venv + pyarrow(读 KaniTTS parquet)本文件为原始资料汇编,未做结论提炼;所有数字均标注了来源与"厂商自评/第三方"属性,未标注者请在引用前回溯到 7 节对应 URL 二次核对。