ElevenLabs vs Cartesia
ElevenLabs Multilingual v2 / Eleven v3 vs Cartesia Sonic 3 / Sonic 3.5 for text-to-speech. ELO: 1179 vs 1203 — edge Cartesia TTFA: 653 ms vs 271 ms — edge Cartesia Langs: 29 vs 42 — edge Cartesia Computed from public benchmarks with dated sources; updated 2026-08-20.
Head to head
| Axis | ElevenLabs Multilingual v2 / Eleven v3 | Cartesia Sonic 3 / Sonic 3.5 | Edge |
|---|---|---|---|
| Price | $100 per 1M chars | $39 per 1M chars* | — |
| ELO | 1179 | 1203 | Cartesia |
| TTFA | 653 ms | 271 ms | Cartesia |
| Langs | 29 | 42 | Cartesia |
| streaming | ✓ | ✓ | — |
| cloning | ✓ | ✓ | — |
* token-/credit-priced — the headline understates real per-unit cost, so no price edge is awarded.
Strengths & caveats
ElevenLabsLargest production adoption and most mature voice-library/cloning ecosystem, with both instant and professional cloning; Eleven v3 is the strongest ElevenLabs model on the Artificial Analysis Speech Arena at 1179 ELO (rank 12 of 99). Most expensive per-character of the comparable group ($100/1M chars at $0.10/1K, charged for both Eleven v3 and Multilingual v2); the $0.05/1K Flash/Turbo and v3 Conversational tiers are cheaper but different models. Independent Coval TTFA is slow for the flagship tiers: 653ms P50 for Eleven v3 and 1262ms for Multilingual v2, against 212ms for the cheaper Flash v2.5 — the 264ms previously listed here was a Turbo v2.5 figure, not a flagship one. Language counts differ per model on the pricing page (29 for Multilingual v2, 32 for Flash/Turbo, 70+ for Eleven v3); the 29 shown is the Multilingual v2 figure. Multilingual v2 ranks only 33rd on the Arena (1104 ELO).CartesiaLowest independently-measured time-to-first-audio in this table (271ms P50 for Sonic 3.5 in the Coval benchmark), and Cartesia's newer Sonic 3.6 — beta, API-only, released 2026-08-18 — currently leads the entire Artificial Analysis Speech Arena at 1285 ELO. Price is credit-/plan-based and the pricing page now states allowances in minutes of audio rather than characters, publishing no per-character rate at all; the ~$39/1M figure is derived from the $49/month Startup tier, so vergelijkbaar=false. Vendor latency claims (~90ms, 'sub-90ms' for 3.6) are roughly three times better than the measured 271ms P50. Sonic 3.5 slipped from about 4th to 7th on the Arena at an unchanged 1203 ELO; Sonic 3.6's much higher score is deliberately not used for this row, because 3.6 is still beta while the docs list 3.5 as stable and priced. The 42-language count carries over from the previous snapshot and was not re-verified in this refresh.
Sources
- SpeechifyAI — API pricing (TTS tiers)2026-08-20
- Artificial Analysis — Text to Speech Leaderboard (Speech Arena provider-voice, blind-vote ELO, 99 models)2026-08-20
- Coval — TTS benchmark data (TTFA P50/P95 and WER, rolling 7-day window on a pinned dataset), mirrored with methodology notes by Openbenchmarks; pulled 2026-08-20 06:30 UTC2026-08-20
- Coval — Best TTS Providers 2026 guide (successor to the 2026-05-04 benchmark post; now carries vendor-claimed latencies only, no benchmark table)2026-08-20
- ElevenLabs — API Pricing2026-08-20
- Cartesia — Pricing2026-08-20
- Google — Gemini Developer API Pricing2026-08-20
- OpenAI — API Pricing (developers.openai.com; the former openai.com/api/pricing/ URL now returns 403)2026-08-20
- MiniMax — Product Pricing (API docs)2026-06-22
- Deepgram — Pricing (Aura-2 and Flux TTS)2026-08-20
- Deepgram — TTS model docs (Aura-2 language list)2026-08-20
- MarkTechPost — Cartesia ships Sonic-3.6 (launch date, beta/API-only status)2026-08-20