Deepgram vs Google
Deepgram Nova-3 vs Google Gemini 3 Flash for transcription. WER: 5.2% vs 2.9% — edge Google Computed from public benchmarks with dated sources; updated 2026-08-20.
Head to head
| Axis | Deepgram Nova-3 | Google Gemini 3 Flash | Edge |
|---|---|---|---|
| Price | $4.3 per 1000 min | $1.92 per 1000 min* | — |
| WER | 5.2% | 2.9% | |
| Langs | 50 | — | — |
| Latency | 300 ms | — | — |
| diarization | ✓ | — | — |
| timestamps | ✓ | — | — |
| vocab | ✓ | — | — |
* token-/credit-priced — the headline understates real per-unit cost, so no price edge is awarded.
Strengths & caveats
DeepgramStreaming-focused with per-second billing and full diarization, timestamps and keyterm support; markets sub-300 ms latency. Lowest accuracy of the group on the leaderboard (5.2% WER); the latency figure is self-reported, not an independent benchmark. Deepgram now positions a separate, dearer Flux model ($0.0077/min list) as its ultra-low-latency option for voice agents.GoogleStrong accuracy (2.9% WER) and broad multilingual coverage as part of a multimodal model. Speech-to-text is a feature of a general LLM, not a dedicated transcription API: no native diarization, word-timestamps or custom vocabulary. The $1.92/1000 min is input-audio tokens only (output transcript billed separately), so it is not directly comparable to per-minute STT pricing.
Sources
- Artificial Analysis — Speech to Text2026-08-20
- Open ASR Leaderboard2026-06-19