Alibaba tops the leaderboard
Qwen-Audio-3.0-TTS-Plus takes #1 on the Artificial Analysis Speech Arena, ahead of Speechify Simba and Gemini 3.1 Flash TTS.
Real benchmarks. Transparent pricing. Guided recommendations.
20+ TTS & STT providers, updated for July 2026.
Independent Benchmarks
Quality, speed, features, and price efficiency across all active providers — sorted by composite value score.
Stacked analysis of Quality, Speed, Features, and Price efficiency. Higher total area = better overall value. Sorted by composite score.
| Provider | Quality (out of 5) | Speed (out of 5) | Features (out of 5) | Price Score (out of 5) | Total (out of 20) |
|---|---|---|---|---|---|
| AssemblyAI | 4.5 | 4 | 5 | 5 | 18.5 |
| Deepgram | 4.5 | 5 | 4 | 4 | 17.5 |
| Azure AI Speech | 4.5 | 4 | 5 | 4 | 17.5 |
| Chatterbox | 4.5 | 4 | 4 | 5 | 17.5 |
| Qwen3-TTS | 4.5 | 4 | 4 | 5 | 17.5 |
| Inworld AI | 5 | 5 | 5 | 2 | 17 |
| Google Gemini 3.1 Flash TTS | 5 | 4 | 5 | 3 | 17 |
| xAI Grok TTS | 4 | 4 | 4 | 5 | 17 |
| Cartesia | 4.5 | 5 | 4 | 3 | 16.5 |
| Hume AI | 4.5 | 4 | 4 | 4 | 16.5 |
| Fish Audio S2 Pro | 4.5 | 4 | 4 | 4 | 16.5 |
| Kokoro | 4.2 | 4 | 3 | 5 | 16.2 |
Scores are composite ratings (1–5 per dimension) compiled from Artificial Analysis, HuggingFace TTS Arena, and official benchmarks. Price Score: 5 = cheapest. Open-source models at $0 self-hosted.
Recommendation Engine
Answer 3 questions to get a personalised recommendation rooted in July 2026 benchmarks.
Step 1 of 3
Market Updates — July 2026
A new leaderboard leader, a sub-150ms transcript API, and open-source models still beating commercial leaders.
Qwen-Audio-3.0-TTS-Plus takes #1 on the Artificial Analysis Speech Arena, ahead of Speechify Simba and Gemini 3.1 Flash TTS.
One HTTP POST returns a finished Universal-3.5 Pro transcript in ~134ms median — no polling or WebSocket.
Natural-language style control via 200+ audio tags, 70+ languages, SynthID watermarking on every output.
$130M Series C (Jan 13) funds the OfOne acquisition; Flux Multilingual (Apr 29) adds sub-300ms turn detection.
Chatterbox-Turbo (MIT) preferred by 65.3% of evaluators vs. ElevenLabs' 24.5% in blind tests.
Terminated Dec 31, 2025 after Meta acquisition; domain no longer resolves. Migrate to ElevenLabs or Chatterbox.
Provider Comparison
Benchmarks, pricing, ELO scores, WER, features, and compliance across every active TTS & STT provider.
| Provider | Gradium BOTH | Inworld AI TTS | ElevenLabs BOTH | Deepgram BOTH | AssemblyAI STT | Cartesia TTS | Alibaba Qwen-Audio-3.0-TTS-PlusNEW TTS | Speechify Simba 3.2NEW TTS | Google Gemini 3.1 Flash TTS TTS | Mistral Voxtral TTS | Microsoft MAINEW BOTH | xAI Grok TTS BOTH | OpenAI BOTH | Azure AI Speech BOTH | Google Cloud Speech BOTH | Hume AI TTS | LeanVox TTS | Kokoro v1.0OPEN SOURCE TTS | ChatterboxOPEN SOURCE TTS | Qwen3-TTSOPEN SOURCE TTS | Fish Audio S2 ProOPEN SOURCE TTS | Moonshine (Useful Sensors)OPEN SOURCE STT | PlayHTDISCONTINUED TTS |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Pricing | Monthly credit plans. XS: $13/month for 225,000 shared credits. TTS: 1 credit/character; STT: 3 credits/second. Free plan available for non-commercial use. Excluded from the usage-only calculator because it does not model shared credits or subscription minimums. Provider pricing | TTS-1.5 Max: $50/1M chars standard ($10/1M on legacy Founder tier) TTS-1.5 Mini: $25/1M chars standard ($5/1M on legacy Founder tier) | Starter: $6/mo for 30k chars Creator: $22/mo for 121k chars Scale API: ~$165/1M chars (Scale) | Nova-3 STT: $0.0043–$0.0077/min (per-second billing) Voice Agent API: ~$0.075/min (STT+LLM+TTS) | Universal-2: $0.0025/min ($0.15/hr) — 99 languages Universal-3.5 Pro: $0.0035/min ($0.21/hr async) — prompt-based customization Sync / Realtime API: $0.0075/min ($0.45/hr) — ~134ms median (Jul 14, 2026) | Pay-as-you-go: $5/100k credits (1 credit/char) | API: $27.60/1M chars | API: $10/1M chars | Flash TTS: ~$18.30/1M chars | API: $16/1M chars | MAI-Transcribe-1.5: ~$0.017/min (Azure Foundry pricing) MAI-Voice-2: $16/1M chars | TTS API: $4.20/1M chars | TTS Standard: $15/1M chars (tts-1) TTS HD: $30/1M chars (tts-1-hd) GPT-4o Transcribe: $0.006/min — free diarization Mini Transcribe: $0.003/min — budget option | Neural TTS (Prebuilt): $16/1M chars Neural HD 2.5: $22/1M chars (down from $30, Mar 2026) STT Standard: $0.017/min (140+ languages) | WaveNet: $4/1M chars (standard) Chirp 3 HD: $30/1M chars (HD) STT V2 Real-time: $0.016/min ($0.004/min Dynamic Batch) | Octave 2: $7.60/1M chars | Standard: $5/1M chars | Self-hosted: Free (compute only) Hosted (DeepInfra): ~$0.65/1M chars hosted | Self-hosted: Free (MIT license) | Self-hosted: Free (Apache 2.0) | API: ~$10/1M chars (commercial API) Research weights: Free — non-commercial only (Fish Audio Research License) | Self-hosted: Free (MIT license) | Service discontinued |
| Quality & ELO | Not scored | Quality Speed ELO 1,198 130ms TTFA | Quality Speed ELO 1,197 75ms TTFA | Quality Speed 5.3% WER | Quality Speed 14.5% WER | Quality Speed ELO 1,208 40ms TTFA | Quality Speed ELO 1,234 | Quality Speed ELO 1,230 | Quality Speed ELO 1,215 | Quality Speed 90ms TTFA | Quality Speed | Quality Speed | Quality Speed 5% WER | Quality Speed | Quality Speed | Quality Speed | Quality Speed | Quality Speed ELO 1,057 | Quality Speed | Quality Speed 1.24% WER | Quality Speed ELO 1,120 | Quality Speed | Quality Speed |
| Key Features |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| Languages | 5+ | 30+ | 74+ | 45+ | 99+ 🌍 | 20+ | 20+ | 20+ | 70+ | 9+ | 43+ | 25+ | 99+ 🌍 | 140+ 🌍 | 75+ | 11+ | 23+ | 9+ | 23+ | 10+ | 80+ | 1+ | N/A |
| Compliance | — | SOC2HIPAA | SOC2HIPAAGDPR | HIPAASOC2 | HIPAASOC2 Type 2ISO 27001:2022PCI DSS v4.0GDPR | — | — | — | GDPRSynthID Provenance | — | Azure ComplianceGDPRHIPAA | — | SOC2GDPR | SOC2HIPAAISO 27001GDPRFedRAMP | SOC2HIPAAISO 27001GDPR | — | — | — | — | — | — | — | — |
| Best For | voice agentcontent creationtranscription | voice agentcontent creationenterprise | content creationnarrationvoice agententerprise | voice agenttranscriptionanalyticsreal time | analyticstranscriptionunderstandingenterprise | voice agentreal time | content creationenterprise | content creationbudget | content creationnarrationenterprise | voice agentbudgetoffline | enterprisetranscriptionaccessibility | voice agentbudgetprototyping | simple appprototypingtranscriptionvoice agent | enterpriseaccessibilityglobal | enterpriseanalyticsaccessibility | voice agentcontent creation | budgetsimple app | budgetaccessibilityoffline | budgetcontent creationvoice agentoffline | budgetcontent creationaccessibility | content creationvoice agent | accessibilityofflinebudget | — |
Data sourced from Artificial Analysis Speech Arena, HuggingFace Open ASR Leaderboard, and official provider documentation. All prices approximate as of July 2026. Benchmark scores may vary by use case.
EU AI Act Article 50 — August 2, 2026 (10 days away). All voice AI systems must disclose AI origin, mark synthetic audio in machine-readable format, and comply with emotion recognition restrictions. The May 2026 Omnibus agreement grants systems already on the market an extension to Dec 2, 2026 for machine-readable marking specifically. California CAITA aligns to the same date. Penalties up to €30M or 7% of global turnover.
Pricing Calculator
Drag the volume slider to compare real costs across all active providers at your scale. Open-source options shown at $0 self-hosted.
Estimated at base pricing tiers.
TTS: ~15,000 chars/hr · STT: 60 mins/hr
Open-source models shown at $0 (self-hosted).
At 10h/month
Community Picks
Ranked from real production deployments, blind tests, and Artificial Analysis benchmarks.