OpenAI vs Deepgram
OpenAI gpt-4o-mini-tts vs Deepgram Aura-2 for text-to-speech. TTFA: 775 ms vs 293 ms — edge Deepgram Langs: 50 vs 7 — edge OpenAI Computed from public benchmarks with dated sources; updated 2026-08-20.
Head to head
| Axis | OpenAI gpt-4o-mini-tts | Deepgram Aura-2 | Edge |
|---|---|---|---|
| Price | $15 per 1M chars* | $30 per 1M chars | — |
| ELO | — | — | — |
| TTFA | 775 ms | 293 ms | Deepgram |
| Langs | 50 | 7 | OpenAI |
| streaming | ✓ | ✓ | — |
| cloning | — | — | — |
* token-/credit-priced — the headline understates real per-unit cost, so no price edge is awarded.
Strengths & caveats
OpenAICheap, developer-friendly, with steerable tone/emotion via a free-text 'instructions' parameter and tight integration into the OpenAI stack; ~11-13 built-in voices. Token-based pricing (unchanged at $0.60/1M text input tokens + $12/1M audio output tokens); the ~$15/1M-chars figure is a community estimate, so vergelijkbaar=false — note OpenAI's older tts-1 is separately listed at a real $15/1M characters, which is easy to confuse with it. Not ranked on the Arena: gpt-4o-mini-tts does not appear among the 99 leaderboard models, where OpenAI is represented only by tts-1, tts-1-hd and GPT-Realtime-2. Independent Coval TTFA is 775ms P50, the slowest in this table apart from ElevenLabs Multilingual v2. The ~50-language claim is vendor-stated and was not re-verified in this refresh. No voice cloning.DeepgramLow, clean per-character price (unchanged at $0.030/1K, $0.027/1K at Growth tier) aimed at voice-agent/IVR workloads with streaming and a large English voice catalog (40+ EN voices). Narrow language coverage, confirmed in the docs at 7 (EN, ES, NL, FR, DE, IT, JA). Independent Coval TTFA improved from 313ms to 293ms P50 — second-best in this table behind Cartesia — but still far off the vendor's '90ms' claim. Aura-2 itself is unranked on the Speech Arena, so naturalness ELO is left null; Deepgram's newer Flux TTS does appear but only at rank 59 (1054 ELO) and costs more at $0.045/1K after a free period ending 2026-09-12. No voice cloning.
Sources
- SpeechifyAI — API pricing (TTS tiers)2026-08-20
- Artificial Analysis — Text to Speech Leaderboard (Speech Arena provider-voice, blind-vote ELO, 99 models)2026-08-20
- Coval — TTS benchmark data (TTFA P50/P95 and WER, rolling 7-day window on a pinned dataset), mirrored with methodology notes by Openbenchmarks; pulled 2026-08-20 06:30 UTC2026-08-20
- Coval — Best TTS Providers 2026 guide (successor to the 2026-05-04 benchmark post; now carries vendor-claimed latencies only, no benchmark table)2026-08-20
- ElevenLabs — API Pricing2026-08-20
- Cartesia — Pricing2026-08-20
- Google — Gemini Developer API Pricing2026-08-20
- OpenAI — API Pricing (developers.openai.com; the former openai.com/api/pricing/ URL now returns 403)2026-08-20
- MiniMax — Product Pricing (API docs)2026-06-22
- Deepgram — Pricing (Aura-2 and Flux TTS)2026-08-20
- Deepgram — TTS model docs (Aura-2 language list)2026-08-20
- MarkTechPost — Cartesia ships Sonic-3.6 (launch date, beta/API-only status)2026-08-20