light estimateLast updated 2026-08-20

Google vs Deepgram

Google Gemini 3.1 Flash TTS vs Deepgram Aura-2 for text-to-speech. Computed from public benchmarks with dated sources; updated 2026-08-20.

Google Gemini 3.1 Flash TTS compared with Deepgram Aura-2 per decision axis
AxisGoogle Gemini 3.1 Flash TTSDeepgram Aura-2Edge
Price$12 per 1M chars*$30 per 1M chars
ELO1212
TTFA293 ms
Langs7
streaming
cloning

* token-/credit-priced — the headline understates real per-unit cost, so no price edge is awarded.

GoogleHighest-scoring row in this table on the Artificial Analysis Speech Arena (1212 ELO, rank 5 of 99), combined with a low effective per-character cost. Token-based pricing (unchanged at $1/1M input text tokens + $20/1M audio output tokens, billed at 25 audio tokens per second); the ~$12/1M-chars figure is a third-party conversion estimate, not a provider per-character rate, so vergelijkbaar=false. Still labelled a Preview model on Google's own pricing page, with more restrictive rate limits. No independent TTFA: Gemini TTS is absent from the Coval benchmark, where Google appears only as chirp-3-hd at 484ms P50. No native voice cloning.DeepgramLow, clean per-character price (unchanged at $0.030/1K, $0.027/1K at Growth tier) aimed at voice-agent/IVR workloads with streaming and a large English voice catalog (40+ EN voices). Narrow language coverage, confirmed in the docs at 7 (EN, ES, NL, FR, DE, IT, JA). Independent Coval TTFA improved from 313ms to 293ms P50 — second-best in this table behind Cartesia — but still far off the vendor's '90ms' claim. Aura-2 itself is unranked on the Speech Arena, so naturalness ELO is left null; Deepgram's newer Flux TTS does appear but only at rank 59 (1054 ELO) and costs more at $0.045/1K after a free period ending 2026-09-12. No voice cloning.