Updated July 2026 — Alibaba tops the board · AssemblyAI Sync API · Gemini 3.1 Flash TTS

The independent voice AI
intelligence radar.

Real benchmarks. Transparent pricing. Guided recommendations.20+ TTS & STT providers, updated for July 2026.

20+
Providers Tracked
$11B
ElevenLabs Valuation
70+
Models on Artificial Analysis
Aug 2
EU AI Act Deadline '26

Independent Benchmarks

Market Leaders at a Glance

Quality, speed, features, and price efficiency across all active providers — sorted by composite value score.

Market Value Composition

Stacked analysis of Quality, Speed, Features, and Price efficiency. Higher total area = better overall value. Sorted by composite score.

Top providers ranked by composite score (Quality + Speed + Features + Price)
ProviderQuality (out of 5)Speed (out of 5)Features (out of 5)Price Score (out of 5)Total (out of 20)
AssemblyAI4.545518.5
Deepgram4.554417.5
Azure AI Speech4.545417.5
Chatterbox4.544517.5
Qwen3-TTS4.544517.5
Inworld AI555217
Google Gemini 3.1 Flash TTS545317
xAI Grok TTS444517
Cartesia4.554316.5
Hume AI4.544416.5
Fish Audio S2 Pro4.544416.5
Kokoro4.243516.2

Scores are composite ratings (1–5 per dimension) compiled from Artificial Analysis, HuggingFace TTS Arena, and official benchmarks. Price Score: 5 = cheapest. Open-source models at $0 self-hosted.

Recommendation Engine

Find Your Perfect Stack

Answer 3 questions to get a personalised recommendation rooted in July 2026 benchmarks.

Step 1 of 3

What are you building?

Market Updates — July 2026

What Changed This Quarter

A new leaderboard leader, a sub-150ms transcript API, and open-source models still beating commercial leaders.

NEW

Alibaba tops the leaderboard

Qwen-Audio-3.0-TTS-Plus takes #1 on the Artificial Analysis Speech Arena, ahead of Speechify Simba and Gemini 3.1 Flash TTS.

NEW

AssemblyAI Sync API (Jul 14)

One HTTP POST returns a finished Universal-3.5 Pro transcript in ~134ms median — no polling or WebSocket.

NEW

Gemini 3.1 Flash TTS (Apr 15)

Natural-language style control via 200+ audio tags, 70+ languages, SynthID watermarking on every output.

MILESTONE

Deepgram — $1.3B

$130M Series C (Jan 13) funds the OfOne acquisition; Flux Multilingual (Apr 29) adds sub-300ms turn detection.

OPEN SOURCE

Chatterbox beats ElevenLabs

Chatterbox-Turbo (MIT) preferred by 65.3% of evaluators vs. ElevenLabs' 24.5% in blind tests.

SHUTDOWN

PlayHT — Discontinued

Terminated Dec 31, 2025 after Meta acquisition; domain no longer resolves. Migrate to ElevenLabs or Chatterbox.

Provider Comparison

Compare All Providers

Benchmarks, pricing, ELO scores, WER, features, and compliance across every active TTS & STT provider.

Filter providers by type
WERELOTTFAMOS— hover for definitions
Provider
Gradium
BOTH
Inworld AI
TTS
ElevenLabs
BOTH
Deepgram
BOTH
AssemblyAI
STT
Cartesia
TTS
Alibaba Qwen-Audio-3.0-TTS-PlusNEW
TTS
Speechify Simba 3.2NEW
TTS
Google Gemini 3.1 Flash TTS
TTS
Mistral Voxtral
TTS
Microsoft MAINEW
BOTH
xAI Grok TTS
BOTH
OpenAI
BOTH
Azure AI Speech
BOTH
Google Cloud Speech
BOTH
Hume AI
TTS
LeanVox
TTS
Kokoro v1.0OPEN SOURCE
TTS
ChatterboxOPEN SOURCE
TTS
Qwen3-TTSOPEN SOURCE
TTS
Fish Audio S2 ProOPEN SOURCE
TTS
Moonshine (Useful Sensors)OPEN SOURCE
STT
PlayHTDISCONTINUED
TTS
Pricing

Monthly credit plans. XS: $13/month for 225,000 shared credits. TTS: 1 credit/character; STT: 3 credits/second. Free plan available for non-commercial use. Excluded from the usage-only calculator because it does not model shared credits or subscription minimums.

Provider pricing
TTS-1.5 Max: $50/1M chars standard ($10/1M on legacy Founder tier)
TTS-1.5 Mini: $25/1M chars standard ($5/1M on legacy Founder tier)
Starter: $6/mo for 30k chars
Creator: $22/mo for 121k chars
Scale API: ~$165/1M chars (Scale)
Nova-3 STT: $0.0043–$0.0077/min (per-second billing)
Voice Agent API: ~$0.075/min (STT+LLM+TTS)
Universal-2: $0.0025/min ($0.15/hr) — 99 languages
Universal-3.5 Pro: $0.0035/min ($0.21/hr async) — prompt-based customization
Sync / Realtime API: $0.0075/min ($0.45/hr) — ~134ms median (Jul 14, 2026)
Pay-as-you-go: $5/100k credits (1 credit/char)
API: $27.60/1M chars
API: $10/1M chars
Flash TTS: ~$18.30/1M chars
API: $16/1M chars
MAI-Transcribe-1.5: ~$0.017/min (Azure Foundry pricing)
MAI-Voice-2: $16/1M chars
TTS API: $4.20/1M chars
TTS Standard: $15/1M chars (tts-1)
TTS HD: $30/1M chars (tts-1-hd)
GPT-4o Transcribe: $0.006/min — free diarization
Mini Transcribe: $0.003/min — budget option
Neural TTS (Prebuilt): $16/1M chars
Neural HD 2.5: $22/1M chars (down from $30, Mar 2026)
STT Standard: $0.017/min (140+ languages)
WaveNet: $4/1M chars (standard)
Chirp 3 HD: $30/1M chars (HD)
STT V2 Real-time: $0.016/min ($0.004/min Dynamic Batch)
Octave 2: $7.60/1M chars
Standard: $5/1M chars
Self-hosted: Free (compute only)
Hosted (DeepInfra): ~$0.65/1M chars hosted
Self-hosted: Free (MIT license)
Self-hosted: Free (Apache 2.0)
API: ~$10/1M chars (commercial API)
Research weights: Free — non-commercial only (Fish Audio Research License)
Self-hosted: Free (MIT license)
Service discontinued
Quality & ELO
Not scored
Quality
Speed
ELO 1,198
130ms TTFA
Quality
Speed
ELO 1,197
75ms TTFA
Quality
Speed
5.3% WER
Quality
Speed
14.5% WER
Quality
Speed
ELO 1,208
40ms TTFA
Quality
Speed
ELO 1,234
Quality
Speed
ELO 1,230
Quality
Speed
ELO 1,215
Quality
Speed
90ms TTFA
Quality
Speed
Quality
Speed
Quality
Speed
5% WER
Quality
Speed
Quality
Speed
Quality
Speed
Quality
Speed
Quality
Speed
ELO 1,057
Quality
Speed
Quality
Speed
1.24% WER
Quality
Speed
ELO 1,120
Quality
Speed
Quality
Speed
Key Features
  • Streaming TTS and STT
  • REST and WebSocket APIs
  • Voice cloning
  • Voice Design from text descriptions
  • Semantic VAD for turn-taking
  • Top-5 TTS Arena ELO
  • Zero-shot Voice Cloning (5–15s)
  • Sub-250ms P90 Latency
  • Realtime TTS-2 (130ms)
  • Domain-specific Pronunciation
  • Healthcare/Finance/Legal
  • Eleven v3 (GA Feb 2)
  • Scribe v2 Realtime STT (sub-150ms, 90+ languages)
  • On-Premise / On-Device
  • Voice Cloning (10,000+ voices)
  • Dubbing & Translation
  • 74 Languages
  • ElevenAgents
  • IBM watsonx Integration
  • Series D: $500M @ $11B (Feb 2026)
  • Nova-3 (5.3% WER)
  • Flux Multilingual (sub-300ms EOT detection)
  • Sub-300ms Streaming
  • Diarization
  • Smart Formatting
  • TTS Speed Controls (0.7–1.5×)
  • Self-hosted Deployment
  • OfOne Acquisition (Restaurant Voice AI)
  • 45+ Languages
  • Per-second Billing
  • Universal-3.5 Pro Streaming
  • Sync API (~134ms median, launched Jul 2026)
  • Context Carryover
  • Prompt-based Domain Customization
  • Three Latency Modes
  • Medical Mode (en/es/de/fr)
  • PII Redaction
  • Voice Agent API ($4.50/hr flat)
  • Audio Intelligence
  • Sonic 3.5 (40ms Turbo TTFB)
  • 3-second Voice Cloning
  • Emotion Control
  • 30+ Featured Voices
  • Dated Immutable Model Snapshots
  • PVC Accent/Gender PATCH API
  • #1 Artificial Analysis Speech Arena
  • Alibaba Cloud Native
  • Highest Human-preference Elo
  • #2 Artificial Analysis Speech Arena
  • Best Elo-to-Price Ratio in Top 5
  • 200+ Audio Tags (Natural-language Style Control)
  • SynthID Watermarking
  • 70+ Languages
  • AI Studio / Vertex AI / Google Vids
  • 4B Parameters
  • 90ms TTFA
  • Smartphone Deployment (3GB RAM)
  • 3-second Voice Adaptation
  • Voxtral Transcribe 2 (ASR)
  • EU Data Sovereignty
  • CC BY-NC 4.0 (open weights)
  • MAI-Voice-2 (72% preferred over v1)
  • MAI-Transcribe-1.5 (43-language WER leader)
  • Emotional Styles & Roles
  • Zero-shot Cloning (5–60s)
  • Copilot/Teams/GitHub Integration
  • Half GPU Usage vs Competitors
  • 5 Voices (Eve, Ara, Leo, Rex, Sal)
  • Speech Tags ([laugh], [sigh], <whisper>)
  • STT — 25 Languages, Batch & Streaming
  • Vercel AI Gateway Integration
  • OpenAI Realtime API Compatible
  • GPT-Realtime-2 (Configurable Reasoning)
  • GPT-Realtime-Translate
  • GPT-Realtime-Whisper (Streaming STT)
  • GPT-4o Transcribe (5% WER)
  • Free Diarization
  • tts-1 / tts-1-hd
  • 99+ Languages (STT)
  • 140+ Languages (TTS)
  • 500+ Neural Voices
  • Neural HD 2.5 (Improved Prosody)
  • Custom Neural Voice
  • Speech Translation
  • Real-time Captions
  • Avatar Video Synthesis
  • Chirp 3 HD
  • WaveNet & Studio Voices
  • Dynamic Batch STT (75% cheaper)
  • Gemini Integration
  • 380+ Voices
  • 75+ Languages
  • Google Translate Integration
  • Octave 2 (sub-200ms)
  • TADA — Open-sourced Mar 2026
  • 5× Faster Inference
  • 700s Long-form Generation
  • Zero Content Hallucinations
  • Emotional Intelligence
  • 11 Languages
  • 23+ Languages
  • Standard Neural Voices
  • REST API
  • 82M Parameters
  • 54 Baked-in Voices
  • MOS 4.2 (highest open-source)
  • CPU / Raspberry Pi Capable
  • 36× Real-time on Free Colab T4
  • Apache 2.0 License
  • 9 Languages
  • MIT License
  • Turbo: 65.3% Preferred over ElevenLabs (24.5%)
  • Chatterbox Turbo (sub-200ms)
  • Chatterbox Multilingual (23+ languages)
  • PerTh Neural Watermarking
  • Paralinguistic Tags [laugh] [cough]
  • Emotion Control
  • Apache 2.0 License
  • 0.77% Chinese WER
  • 1.24% English WER
  • 0.6B & 1.7B Variants
  • Base / CustomVoice / VoiceDesign
  • Default HF Speech-to-Speech TTS
  • 49+ Voice Presets
  • 10 Languages
  • Elo 1,120 (Top Open-weights)
  • Built on Qwen3-4B Backbone
  • Inline Emotion Cues ([whisper], [laugh])
  • 80+ Languages
  • Research License (Non-commercial Free)
  • 245M Parameters (MIT)
  • Matches Whisper Large-v3
  • 1/6 the Size of Whisper
  • Mobile & Embedded Ready
  • CPU Capable
  • Service Discontinued (Dec 31, 2025)
  • Team Absorbed into Meta Superintelligence Labs
  • No Data Migration Path
  • Migrate to: ElevenLabs, Chatterbox, Kokoro
Languages5+ 30+ 74+ 45+ 99+ 🌍20+ 20+ 20+ 70+ 9+ 43+ 25+ 99+ 🌍140+ 🌍75+ 11+ 23+ 9+ 23+ 10+ 80+ 1+ N/A
Compliance
SOC2HIPAA
SOC2HIPAAGDPR
HIPAASOC2
HIPAASOC2 Type 2ISO 27001:2022PCI DSS v4.0GDPR
GDPRSynthID Provenance
Azure ComplianceGDPRHIPAA
SOC2GDPR
SOC2HIPAAISO 27001GDPRFedRAMP
SOC2HIPAAISO 27001GDPR
Best For
voice agentcontent creationtranscription
voice agentcontent creationenterprise
content creationnarrationvoice agententerprise
voice agenttranscriptionanalyticsreal time
analyticstranscriptionunderstandingenterprise
voice agentreal time
content creationenterprise
content creationbudget
content creationnarrationenterprise
voice agentbudgetoffline
enterprisetranscriptionaccessibility
voice agentbudgetprototyping
simple appprototypingtranscriptionvoice agent
enterpriseaccessibilityglobal
enterpriseanalyticsaccessibility
voice agentcontent creation
budgetsimple app
budgetaccessibilityoffline
budgetcontent creationvoice agentoffline
budgetcontent creationaccessibility
content creationvoice agent
accessibilityofflinebudget

Data sourced from Artificial Analysis Speech Arena, HuggingFace Open ASR Leaderboard, and official provider documentation. All prices approximate as of July 2026. Benchmark scores may vary by use case.

EU AI Act Article 50 — August 2, 2026 (10 days away). All voice AI systems must disclose AI origin, mark synthetic audio in machine-readable format, and comply with emotion recognition restrictions. The May 2026 Omnibus agreement grants systems already on the market an extension to Dec 2, 2026 for machine-readable marking specifically. California CAITA aligns to the same date. Penalties up to €30M or 7% of global turnover.

Pricing Calculator

Estimate Your Monthly Cost

Drag the volume slider to compare real costs across all active providers at your scale. Open-source options shown at $0 self-hosted.

Technology
10 hours

Estimated at base pricing tiers.

TTS: ~15,000 chars/hr · STT: 60 mins/hr

Open-source models shown at $0 (self-hosted).

At 10h/month

Cheapest paid:$0.60Google Cloud Speech
Most expensive:$30.00ElevenLabs

Monthly Cost Comparison

Commercial APIOpen Source (self-hosted $0)

Community Picks

Best Provider by Use Case

Ranked from real production deployments, blind tests, and Artificial Analysis benchmarks.

Real-time Voice Agent

  1. 1.Cartesia Sonic 3.5 — 40ms TTFA
  2. 2.Inworld Realtime TTS-2 — 130ms
  3. 3.ElevenLabs Flash v2.5 — ~75ms

Highest Quality TTS

  1. 1.Alibaba Qwen-Audio-3.0-TTS-Plus (ELO 1,234)
  2. 2.Speechify Simba 3.2 (ELO 1,230)
  3. 3.Gemini 3.1 Flash TTS (ELO 1,215)
💰

Best Value STT

  1. 1.AssemblyAI Universal-2 — $0.0025/min
  2. 2.Deepgram Nova-3 — $0.0043/min
  3. 3.OpenAI Mini Transcribe — $0.003/min
🎤

Voice Cloning

  1. 1.ElevenLabs — professional grade
  2. 2.Cartesia — 3-second cloning
  3. 3.Chatterbox — free, MIT licensed
📱

Edge / Offline TTS

  1. 1.Kokoro 82M — CPU capable, Apache 2.0
  2. 2.Chatterbox Turbo — sub-200ms
  3. 3.Mistral Voxtral — 3 GB RAM
🔒

Edge / Offline STT

  1. 1.Moonshine 245M — MIT, = Whisper v3
  2. 2.Whisper.cpp — 38K+ GitHub stars
  3. 3.NVIDIA Canary Qwen — 5.63% WER