Back to overview

Provider Comparison Matrix

Benchmarks, pricing models, ELO scores, word error rates, feature sets, and compliance certifications across every active TTS & STT provider. Updated July 2026.

Filter providers by type
WERELOTTFAMOS— hover for definitions
Provider
Gradium
BOTH
Inworld AI
TTS
ElevenLabs
BOTH
Deepgram
BOTH
AssemblyAI
STT
Cartesia
TTS
Alibaba Qwen-Audio-3.0-TTS-PlusNEW
TTS
Speechify Simba 3.2NEW
TTS
Google Gemini 3.1 Flash TTS
TTS
Mistral Voxtral
TTS
Microsoft MAINEW
BOTH
xAI Grok TTS
BOTH
OpenAI
BOTH
Azure AI Speech
BOTH
Google Cloud Speech
BOTH
Hume AI
TTS
LeanVox
TTS
Kokoro v1.0OPEN SOURCE
TTS
ChatterboxOPEN SOURCE
TTS
Qwen3-TTSOPEN SOURCE
TTS
Fish Audio S2 ProOPEN SOURCE
TTS
Moonshine (Useful Sensors)OPEN SOURCE
STT
PlayHTDISCONTINUED
TTS
Pricing

Monthly credit plans. XS: $13/month for 225,000 shared credits. TTS: 1 credit/character; STT: 3 credits/second. Free plan available for non-commercial use. Excluded from the usage-only calculator because it does not model shared credits or subscription minimums.

Provider pricing
TTS-1.5 Max: $50/1M chars standard ($10/1M on legacy Founder tier)
TTS-1.5 Mini: $25/1M chars standard ($5/1M on legacy Founder tier)
Starter: $6/mo for 30k chars
Creator: $22/mo for 121k chars
Scale API: ~$165/1M chars (Scale)
Nova-3 STT: $0.0043–$0.0077/min (per-second billing)
Voice Agent API: ~$0.075/min (STT+LLM+TTS)
Universal-2: $0.0025/min ($0.15/hr) — 99 languages
Universal-3.5 Pro: $0.0035/min ($0.21/hr async) — prompt-based customization
Sync / Realtime API: $0.0075/min ($0.45/hr) — ~134ms median (Jul 14, 2026)
Pay-as-you-go: $5/100k credits (1 credit/char)
API: $27.60/1M chars
API: $10/1M chars
Flash TTS: ~$18.30/1M chars
API: $16/1M chars
MAI-Transcribe-1.5: ~$0.017/min (Azure Foundry pricing)
MAI-Voice-2: $16/1M chars
TTS API: $4.20/1M chars
TTS Standard: $15/1M chars (tts-1)
TTS HD: $30/1M chars (tts-1-hd)
GPT-4o Transcribe: $0.006/min — free diarization
Mini Transcribe: $0.003/min — budget option
Neural TTS (Prebuilt): $16/1M chars
Neural HD 2.5: $22/1M chars (down from $30, Mar 2026)
STT Standard: $0.017/min (140+ languages)
WaveNet: $4/1M chars (standard)
Chirp 3 HD: $30/1M chars (HD)
STT V2 Real-time: $0.016/min ($0.004/min Dynamic Batch)
Octave 2: $7.60/1M chars
Standard: $5/1M chars
Self-hosted: Free (compute only)
Hosted (DeepInfra): ~$0.65/1M chars hosted
Self-hosted: Free (MIT license)
Self-hosted: Free (Apache 2.0)
API: ~$10/1M chars (commercial API)
Research weights: Free — non-commercial only (Fish Audio Research License)
Self-hosted: Free (MIT license)
Service discontinued
Quality & ELO
Not scored
Quality
Speed
ELO 1,198
130ms TTFA
Quality
Speed
ELO 1,197
75ms TTFA
Quality
Speed
5.3% WER
Quality
Speed
14.5% WER
Quality
Speed
ELO 1,208
40ms TTFA
Quality
Speed
ELO 1,234
Quality
Speed
ELO 1,230
Quality
Speed
ELO 1,215
Quality
Speed
90ms TTFA
Quality
Speed
Quality
Speed
Quality
Speed
5% WER
Quality
Speed
Quality
Speed
Quality
Speed
Quality
Speed
Quality
Speed
ELO 1,057
Quality
Speed
Quality
Speed
1.24% WER
Quality
Speed
ELO 1,120
Quality
Speed
Quality
Speed
Key Features
  • Streaming TTS and STT
  • REST and WebSocket APIs
  • Voice cloning
  • Voice Design from text descriptions
  • Semantic VAD for turn-taking
  • Top-5 TTS Arena ELO
  • Zero-shot Voice Cloning (5–15s)
  • Sub-250ms P90 Latency
  • Realtime TTS-2 (130ms)
  • Domain-specific Pronunciation
  • Healthcare/Finance/Legal
  • Eleven v3 (GA Feb 2)
  • Scribe v2 Realtime STT (sub-150ms, 90+ languages)
  • On-Premise / On-Device
  • Voice Cloning (10,000+ voices)
  • Dubbing & Translation
  • 74 Languages
  • ElevenAgents
  • IBM watsonx Integration
  • Series D: $500M @ $11B (Feb 2026)
  • Nova-3 (5.3% WER)
  • Flux Multilingual (sub-300ms EOT detection)
  • Sub-300ms Streaming
  • Diarization
  • Smart Formatting
  • TTS Speed Controls (0.7–1.5×)
  • Self-hosted Deployment
  • OfOne Acquisition (Restaurant Voice AI)
  • 45+ Languages
  • Per-second Billing
  • Universal-3.5 Pro Streaming
  • Sync API (~134ms median, launched Jul 2026)
  • Context Carryover
  • Prompt-based Domain Customization
  • Three Latency Modes
  • Medical Mode (en/es/de/fr)
  • PII Redaction
  • Voice Agent API ($4.50/hr flat)
  • Audio Intelligence
  • Sonic 3.5 (40ms Turbo TTFB)
  • 3-second Voice Cloning
  • Emotion Control
  • 30+ Featured Voices
  • Dated Immutable Model Snapshots
  • PVC Accent/Gender PATCH API
  • #1 Artificial Analysis Speech Arena
  • Alibaba Cloud Native
  • Highest Human-preference Elo
  • #2 Artificial Analysis Speech Arena
  • Best Elo-to-Price Ratio in Top 5
  • 200+ Audio Tags (Natural-language Style Control)
  • SynthID Watermarking
  • 70+ Languages
  • AI Studio / Vertex AI / Google Vids
  • 4B Parameters
  • 90ms TTFA
  • Smartphone Deployment (3GB RAM)
  • 3-second Voice Adaptation
  • Voxtral Transcribe 2 (ASR)
  • EU Data Sovereignty
  • CC BY-NC 4.0 (open weights)
  • MAI-Voice-2 (72% preferred over v1)
  • MAI-Transcribe-1.5 (43-language WER leader)
  • Emotional Styles & Roles
  • Zero-shot Cloning (5–60s)
  • Copilot/Teams/GitHub Integration
  • Half GPU Usage vs Competitors
  • 5 Voices (Eve, Ara, Leo, Rex, Sal)
  • Speech Tags ([laugh], [sigh], <whisper>)
  • STT — 25 Languages, Batch & Streaming
  • Vercel AI Gateway Integration
  • OpenAI Realtime API Compatible
  • GPT-Realtime-2 (Configurable Reasoning)
  • GPT-Realtime-Translate
  • GPT-Realtime-Whisper (Streaming STT)
  • GPT-4o Transcribe (5% WER)
  • Free Diarization
  • tts-1 / tts-1-hd
  • 99+ Languages (STT)
  • 140+ Languages (TTS)
  • 500+ Neural Voices
  • Neural HD 2.5 (Improved Prosody)
  • Custom Neural Voice
  • Speech Translation
  • Real-time Captions
  • Avatar Video Synthesis
  • Chirp 3 HD
  • WaveNet & Studio Voices
  • Dynamic Batch STT (75% cheaper)
  • Gemini Integration
  • 380+ Voices
  • 75+ Languages
  • Google Translate Integration
  • Octave 2 (sub-200ms)
  • TADA — Open-sourced Mar 2026
  • 5× Faster Inference
  • 700s Long-form Generation
  • Zero Content Hallucinations
  • Emotional Intelligence
  • 11 Languages
  • 23+ Languages
  • Standard Neural Voices
  • REST API
  • 82M Parameters
  • 54 Baked-in Voices
  • MOS 4.2 (highest open-source)
  • CPU / Raspberry Pi Capable
  • 36× Real-time on Free Colab T4
  • Apache 2.0 License
  • 9 Languages
  • MIT License
  • Turbo: 65.3% Preferred over ElevenLabs (24.5%)
  • Chatterbox Turbo (sub-200ms)
  • Chatterbox Multilingual (23+ languages)
  • PerTh Neural Watermarking
  • Paralinguistic Tags [laugh] [cough]
  • Emotion Control
  • Apache 2.0 License
  • 0.77% Chinese WER
  • 1.24% English WER
  • 0.6B & 1.7B Variants
  • Base / CustomVoice / VoiceDesign
  • Default HF Speech-to-Speech TTS
  • 49+ Voice Presets
  • 10 Languages
  • Elo 1,120 (Top Open-weights)
  • Built on Qwen3-4B Backbone
  • Inline Emotion Cues ([whisper], [laugh])
  • 80+ Languages
  • Research License (Non-commercial Free)
  • 245M Parameters (MIT)
  • Matches Whisper Large-v3
  • 1/6 the Size of Whisper
  • Mobile & Embedded Ready
  • CPU Capable
  • Service Discontinued (Dec 31, 2025)
  • Team Absorbed into Meta Superintelligence Labs
  • No Data Migration Path
  • Migrate to: ElevenLabs, Chatterbox, Kokoro
Languages5+ 30+ 74+ 45+ 99+ 🌍20+ 20+ 20+ 70+ 9+ 43+ 25+ 99+ 🌍140+ 🌍75+ 11+ 23+ 9+ 23+ 10+ 80+ 1+ N/A
Compliance
SOC2HIPAA
SOC2HIPAAGDPR
HIPAASOC2
HIPAASOC2 Type 2ISO 27001:2022PCI DSS v4.0GDPR
GDPRSynthID Provenance
Azure ComplianceGDPRHIPAA
SOC2GDPR
SOC2HIPAAISO 27001GDPRFedRAMP
SOC2HIPAAISO 27001GDPR
Best For
voice agentcontent creationtranscription
voice agentcontent creationenterprise
content creationnarrationvoice agententerprise
voice agenttranscriptionanalyticsreal time
analyticstranscriptionunderstandingenterprise
voice agentreal time
content creationenterprise
content creationbudget
content creationnarrationenterprise
voice agentbudgetoffline
enterprisetranscriptionaccessibility
voice agentbudgetprototyping
simple appprototypingtranscriptionvoice agent
enterpriseaccessibilityglobal
enterpriseanalyticsaccessibility
voice agentcontent creation
budgetsimple app
budgetaccessibilityoffline
budgetcontent creationvoice agentoffline
budgetcontent creationaccessibility
content creationvoice agent
accessibilityofflinebudget

Data sourced from Artificial Analysis Speech Arena, HuggingFace Open ASR Leaderboard, and official provider documentation. All prices approximate as of July 2026. Benchmark scores may vary by use case.