Best Voice AI Models in August 2026: STT, TTS, and Voice Agent Stack
August 2026 voice AI picks: Cartesia Sonic-3.6 tops both speech arenas, Deepgram Flux TTS ships sub-100ms, free through September, plus the STT and agent stack.
Table of Contents
You are shipping a voice agent, and the ground moved in August. Two conversation-native streaming text-to-speech models landed five days apart, both answering in under 100ms on vendor timing, and one of them took the top spot on both independent speech arenas within a week.
Streaming transcription runs turn-by-turn under half a second, a fast LLM closes the loop, and the hard part is choosing among components that all work. This guide picks the pieces of a production voice agent and the latency budget they must fit.

TL;DR: Best voice AI per layer, August 2026
| Layer | Best pick | Why | Pricing |
|---|---|---|---|
| TTS quality (independent) | Cartesia Sonic-3.6 | #1 on both Artificial Analysis speech arenas (Aug 19), sub-90ms vendor TTFA | ~$49/1M chars; agents $0.06/min |
| TTS conversation-native | Deepgram Flux TTS | Vendor “as low as 80ms”, model-integrated turns; free through Sept 12 | $0.045/1k chars |
| TTS expressive | Eleven v3 | 70+ languages, audio tags, not real-time | $0.10/1k chars |
| TTS low-latency (ElevenLabs) | ElevenLabs Flash v2.5 | ~75ms vendor inference, Turbo now deprecated, 40k-char limit | PAYG char |
| TTS enterprise | Deepgram Aura-2 | Sub-200ms vendor steady-state | $0.030/1k chars |
| STT streaming (agents) | AssemblyAI Universal-3.5 Pro Realtime | 4.1% WER on AA-WER Streaming, ~0.4s to first final | $0.45/hr ($7.50/1k min) |
| STT accuracy (batch) | AssemblyAI Universal-3.5 Pro | 7.0% aggregate WER, promptable transcription | $0.21/hr batch |
| STT accuracy-speed | Microsoft MAI-Transcribe-1.5 | 2.4% WER, #3 independent, leads the accuracy-speed frontier | $6/1k min (Foundry) |
| STT cost / broad languages | Deepgram Nova-3 | Added Afrikaans and Georgian in August | $0.0048/min |
| STT turn-taking | Deepgram Flux (STT) | Model-integrated end-of-turn under 400ms | $0.0065/min EN |
| STT (ElevenLabs stacks) | ElevenLabs Scribe v2 Realtime | Aug 3: entity detection plus turbo and lite tiers | PAYG |
| Voice agent (managed) | Retell AI | Single-number billing, HIPAA on the Enterprise tier | from $0.07+/min |
| Voice agent at scale (BYO) | Vapi | $0.05/min platform fee, bring your own stack | ~$0.23-0.33/min all-in |
| Open-source orchestration | LiveKit Agents / Pipecat | Apache-2.0 / BSD-2-Clause, self-host free | infra + providers |
| Speech-to-speech | OpenAI Realtime (gpt-realtime-2.1) | Single-model audio loop (July release) | per audio token |
| LLM brain | See the August LLM guide | Pick a fast model to keep the loop tight | cross-reference |
If you only read one row: Deepgram Nova-3 plus Flux for STT, a fast LLM brain such as Gemini 3.7 Flash or Grok 4.6 (see the August LLM guide), Cartesia Sonic-3.6 or Deepgram Flux TTS for speech, and Retell or Vapi to orchestrate.
That stack lands inside the practical round-trip budget and runs at production scale today.

The story of voice AI in August 2026
Voice AI spent early 2026 turning research demos into datasheet-grade components, and August was the month text-to-speech caught up to the rest of the stack on latency. Two conversation-native streaming models launched inside a five-day window, both quoting sub-100ms time-to-first-audio, and the independent leaderboard confirmed one of them within days.
The engineering question is no longer whether a layer works. It is which pieces compose for your accents, languages, and cost ceiling.
The first launch was Deepgram Flux TTS on August 12, a conversation-aware model that quotes latency “as low as 80ms” even under load and handles turn-taking inside the model rather than leaving it to the application.
It lists at $0.045 per 1,000 characters after trial and is free through September 12, 2026, with 45 concurrent streams globally and 5 in the EU and Australia, moving to standard pricing on September 13.
Flux TTS pairs with Deepgram’s Flux STT, so a team can run a matched conversational STT and TTS pair from one vendor.
Five days later, Cartesia Sonic-3.6 shipped in beta on August 17, adding Odia and Urdu to reach 44 languages at a vendor time-to-first-audio under 90ms.
On August 18 to 19 it ranked first on both Artificial Analysis speech arenas, leading Provider Voice at roughly 1,283 Elo and Controlled Voice at roughly 1,123 Elo.
A vendor launch paired with an independent benchmark inside the same week is the cleanest signal the category produced in August, and it makes sub-100ms streaming TTS table stakes rather than a headline feature.
The beta pricing and exact general-availability date are still settling, so confirm the final list price on Cartesia’s own page.
The dated August news on transcription was ElevenLabs, whose Scribe v2 Realtime update on August 3 added entity detection, new scribe_v2_realtime_turbo and _lite IDs, and a secondary_languages parameter. Deepgram Nova-3 added Afrikaans and Georgian.
Microsoft MAI-Transcribe-1.5 still leads the accuracy-speed frontier at 2.4% word error rate, third on the raw independent board and up to five times faster than its peers, but it launched at Build on June 2, so it belongs on any shortlist as a standing pick rather than as August news.
The business story stayed quiet. No August funding, acquisition, or pricing move surfaced for Vapi, Retell, LiveKit, or Pipecat, and the reported ElevenLabs tender offer near a $22 billion valuation remains pending per July reporting, with its Series D closed back on February 4, 2026 at $11 billion led by Sequoia.
The round-trip target still anchors on ITU-T G.114, which sets one-way mouth-to-ear delay at 150ms preferred and 400ms tolerable. Hitting that budget in August 2026 is a matter of picking components that compose, then measuring your real p50 and p95 on your own traffic.
Best speech-to-text (STT) models in August 2026
STT split cleanly this month by job. One model leads raw accuracy per unit of speed, one leads streaming for agents, one leads cost and coverage, and one solves turn-taking. Batch word error rate does not predict streaming behavior, so match the benchmark to how you will actually run the model.
Microsoft MAI-Transcribe-1.5. The accuracy-speed leader
The first-party cloud STT that leads the accuracy-speed frontier, and a standing pick rather than an August launch, since it shipped at Build on June 2. MAI-Transcribe-1.5 posts 2.4% word error rate on the independent board, third overall on raw accuracy, while running up to five times faster than Gemini 3.1 Flash, Scribe v2, and gpt-4o-transcribe.
Specs:
- WER: 2.4% on Artificial Analysis (independent); #3 on raw WER, first on the accuracy-speed Pareto frontier
- Throughput: up to 5x faster than comparable models
- Availability: Microsoft Foundry (launched June 2 at Build, not an August event)
- Pricing: $6 per 1,000 minutes via Microsoft Foundry
Best for: Teams on Azure or Foundry that want the best accuracy per dollar per unit of speed, batch and near-real-time transcription at scale.
Skip if: You are not on Azure or Foundry, or you need model-native streaming turn-taking today (pair Nova-3 with Flux).
AssemblyAI Universal-3.5 Pro Realtime. Streaming with context
The pick when an agent needs turn-by-turn transcription without dropping context or reconnecting. Universal-3.5 Pro Realtime measures 4.1% word error rate on the AA-WER Streaming benchmark and returns a first final in roughly 0.4 seconds, and the batch sibling remains a high-accuracy async option.
Specs:
- WER: 4.1% on AA-WER Streaming (independent) for the Realtime model; batch Universal-3.5 Pro at 7.0% aggregate WER
- Latency: ~0.4s to first final, per-second billing
- Pricing: $0.45/hr streaming ($7.50 per 1,000 minutes); $0.21/hr batch
- Promptable transcription and structure on the streaming transcript
Best for: Multi-turn agents that need context carried across the call, multi-speaker calls, and domain jargon generic models mistranscribe.
Skip if: You need sub-$0.30 per hour transcription (use Deepgram Nova-3) or pure real-time dictation.
Deepgram Nova-3 and Flux STT. Cost, coverage, and turn-taking
The cost and coverage default, now with wider language support. Nova-3 added Afrikaans and Georgian in August and stays the streaming price leader, and Deepgram Flux STT adds model-integrated end-of-turn detection for agents that need to know when a speaker has actually finished.
Specs:
- Nova-3: broad language coverage plus August additions (Afrikaans, Georgian), $0.0048/min streaming
- Flux STT: model-integrated end-of-turn under 400ms, $0.0065/min EN
- Both tuned for the sub-300ms budgets voice agents live in
- Pairs natively with Deepgram Flux TTS for a matched conversational stack
Best for: Production agents where cost and language breadth are the binding constraints, and any turn-based agent that needs semantic end-of-turn rather than a fixed pause timer.
Skip if: You want the lowest independently measured accuracy figure (use MAI-Transcribe-1.5 or AssemblyAI) or pure dictation with no turn-taking.
ElevenLabs Scribe v2 Realtime. The August STT update
The dated August STT news for ElevenLabs stacks. The August 3 changelog added entity detection to Scribe v2 Realtime, new turbo and lite model IDs, and a secondary-languages parameter, keeping ElevenLabs transcription inside a live-agent latency window.
Specs:
- August 3 update: entity detection,
scribe_v2_realtime_turboand_liteIDs,secondary_languages - Latency: realtime, tuned for live agents
- Best fit inside an existing ElevenLabs voice stack
- Pricing: pay as you go
Best for: Teams already standardized on ElevenLabs voices that want realtime transcription with structured entity output in the same stack.
Skip if: You need a published, independently measured word error rate to budget against (use MAI-Transcribe-1.5 or AssemblyAI).
Best text-to-speech (TTS) models in August 2026
August was the biggest TTS month of the year. Two conversation-native streaming models launched inside five days, both quoting sub-100ms time-to-first-audio, and one topped the independent quality board. Our text-to-speech providers guide tracks the broader field these two now lead.
One caution up front: every vendor time-to-first-audio below is model latency measured in isolation, not end-to-end, and independent p50 streaming latency runs higher.

Cartesia Sonic-3.6. The quality and latency leader
The structural pick when both naturalness and round-trip latency matter. Sonic-3.6 shipped in beta on August 17 with a vendor time-to-first-audio under 90ms, then ranked first on both Artificial Analysis speech arenas on August 18 to 19, so it leads quality and latency at the same time.
Specs:
- Quality: #1 on Provider Voice (~1,283 Elo) and Controlled Voice (~1,123 Elo) on Artificial Analysis (Aug 18 to 19)
- TTFA: sub-90ms vendor model latency, not end-to-end
- Languages: 44 (added Odia and Urdu)
- Pricing: ~$49 per 1M characters (aggregator-normalized), Scale $299/mo, voice agents $0.06/min; beta, confirm final pricing
Best for: The lowest-latency, most natural streaming TTS, telephony where latency drives perceived call quality, real-time interactive apps.
Skip if: You need open weights (none) or a settled general-availability price today (it is still in beta).
Deepgram Flux TTS. Conversation-native TTS
The conversation-aware TTS launched August 12, built to handle turn-taking inside the model rather than at the application layer. Flux TTS quotes latency “as low as 80ms” even under load, and it pairs with Deepgram Flux STT for a matched conversational STT and TTS pair from one vendor.
Specs:
- Latency: “as low as 80ms” vendor model latency, even under load
- Turn handling: model-integrated, conversation-aware
- Pricing: $0.045 per 1,000 characters after trial; free through September 12, 2026 (45 concurrent streams global, 5 EU and Australia); standard from September 13
- Pairs with Deepgram Flux STT for a matched stack
Best for: Enterprise agents that want a single-vendor conversational STT and TTS pair, and teams that can adopt it before the free window closes.
Skip if: You need it free past September 12, or you want a long production track record (it is brand-new).
ElevenLabs Flash v2.5 and Eleven v3. The ElevenLabs paths
The two ElevenLabs lanes. Flash v2.5 is the real-time conversational path near 75ms model inference, and in August ElevenLabs deprecated its Turbo models and named Flash the replacement, with a 40,000-character request limit. Eleven v3 remains the expressive model for non-real-time work.
Specs:
- Flash v2.5: ~75ms vendor model inference, Turbo deprecated in August, 40,000-character request limit
- Eleven v3: 70+ languages, audio tags, expressive; generally available since February 2026, not real-time
- Pricing: Flash v2.5 pay as you go, character-based; Eleven v3 at $0.10 per 1,000 characters
- Voice cloning supported
Best for: Real-time conversational agents that want ElevenLabs voices (Flash v2.5); audiobooks, narration, and character work rendered ahead of time (Eleven v3).
Skip if: You need the lowest possible time-to-first-audio (use Cartesia Sonic-3.6) or real-time output from v3 (use Flash v2.5).
Deepgram Aura-2. Enterprise TTS
The pick when an existing Deepgram enterprise deployment wants first-party TTS end to end. Aura-2 carries a broad voice set at a vendor sub-200ms steady-state latency, which fits teams already running Nova-3 and Flux STT.
Specs:
- Latency: sub-200ms vendor steady-state; higher on independent p50 streaming
- Pricing: $0.030 per 1,000 characters (pay as you go)
- First-party fit with the Deepgram STT stack
Best for: Existing Deepgram enterprise users that want one vendor for STT and TTS.
Skip if: You want the newer conversation-native Flux TTS, or the highest independent quality score (use Sonic-3.6).
Microsoft MAI-Voice-2. The hyperscaler expressive path
The Foundry expressive TTS line, and like MAI-Transcribe a standing pick rather than an August launch, since it reached general availability at Build on June 2. MAI-Voice-2 covers 15 languages with short-sample cloning, and the MAI-Voice-2-Flash preview is the faster, lower-cost variant.
Specs:
- MAI-Voice-2: expressive TTS, 15 languages, short-sample voice cloning (GA June 2 at Build)
- MAI-Voice-2-Flash: preview, “2x speed at 32% lower cost” per vendor
- Availability: Microsoft Foundry
- Pricing: Foundry pricing, not fully re-confirmed
Best for: Azure, Foundry, and Copilot stacks that want first-party expressive TTS.
Skip if: You are not on Azure, or you need a published, independently measured time-to-first-audio to budget against (use Cartesia or ElevenLabs).
Best voice agent platforms in August 2026
If you do not want to wire STT, LLM, TTS, and orchestration yourself, the platforms below ship in days what a custom build ships in quarters. No August pricing, funding, or compliance change surfaced for any of them, so treat the numbers here as the July baseline and confirm the current tier before you build.
Retell AI. The most-teams default
The right default for most production voice-agent teams. Retell bundles HIPAA and a BAA on its Enterprise tier and bills through a single number, which removes the passthrough and compliance overhead that slows a build.
Specs:
- Pricing: from $0.07+ per minute, single-vendor billing
- Builder: no-code visual builder plus developer SDK
- Compliance: HIPAA and BAA bundled (Enterprise tier; re-verify)
Best for: Most production voice agents that want managed HIPAA and a platform that reduces engineering load.
Skip if: You want to bring your own every component (use Vapi) or a fully self-hosted stack (use LiveKit or Pipecat).
Vapi. The scale and BYO pick
The pick when you want to bring your own STT, LLM, and TTS and run at volume. Vapi charges a $0.05 per minute platform fee on top of component passthrough, landing all-in around $0.23 to $0.33 per minute.
Specs:
- Pricing: $0.05/min platform fee plus passthrough; all-in ~$0.23-0.33/min BYOK
- Bring-your-own STT, LLM, and TTS
- Compliance: HIPAA BAA add-on at $2,000/mo (baseline; re-verify)
Best for: Teams that want component-level control and volume scale.
Skip if: You are cost-sensitive at scale or need a cheap BAA (Retell bundles it), or you want a self-hosted stack.
OpenAI Realtime API. Native speech-to-speech
The pick when you want one model for the whole audio loop. The Realtime API runs gpt-realtime-2.1, which released July 6 and cut p95 latency by at least 25% versus the prior model, priced per audio token rather than per minute.
Specs:
- Model: gpt-realtime-2.1 and mini, speech-to-speech (July 6 release, current in August)
- Latency: p95 down at least 25% versus prior
- Pricing: per audio token; mini priced as the prior gpt-realtime-mini
Best for: Prosody-sensitive products and single-vendor OpenAI stacks that value a simpler loop over per-minute predictability.
Skip if: You need per-minute cost control, or separable STT and TTS control.
LiveKit Agents and Pipecat. Open-source orchestration
The open-source route when you want to own the orchestration and still get first-party adapters for every STT and TTS. LiveKit Agents is Apache-2.0 with a managed Cloud, and Pipecat is BSD-2-Clause and fully self-composed.
Specs:
- LiveKit Agents: Apache-2.0, self-host free; Cloud adds $0.01/min agent sessions; signed BAA at Scale ($500/mo, baseline)
- Pipecat: BSD-2-Clause framework, free; plugin ecosystem across Cartesia, Deepgram, ElevenLabs, and OpenAI
- Compliance: LiveKit BAA at Scale; Pipecat is your own deployment boundary
Best for: Teams that want permissive open-source orchestration with production infrastructure (LiveKit) or full control of the orchestration code (Pipecat).
Skip if: You want a fully managed no-code platform (use Retell) or a signed BAA below the Scale tier.
HIPAA and compliance tier
Compliance is the cleanest differentiator across managed platforms, and every signed BAA sits behind an enterprise plan or a paid add-on. These figures carry from July, so confirm the exact tier before you build:
| Platform | HIPAA / signed BAA | Where it lands |
|---|---|---|
| Retell AI | Bundled | HIPAA and BAA included in the flat rate (Enterprise; re-verify) |
| Vapi | $2,000/mo add-on | Platform fee plus passthrough on top |
| LiveKit Agents | Scale tier | Signed BAA at Scale ($500/mo baseline) |
| Pipecat / self-host | Your own deployment | You own the compliance boundary |
Vendor WER versus independent benchmarks
Every STT vendor publishes a best-case word error rate, and those numbers run optimistic against a neutral board. The independent Artificial Analysis index is the cleaner cross-provider read, our speech-to-text APIs guide breaks the providers down on pricing and benchmarks, and batch accuracy does not predict streaming behavior.
On the TTS side, vendor time-to-first-audio is a floor, not what your users hear, so a domain reproduction with your accents, your background noise, and your prompts is the only number that decides the pick.
End-to-end latency budget. The math
Voice agents have a hard latency target. The ITU-T G.114 one-way mouth-to-ear recommendation is 150ms preferred and 400ms tolerable, and that constraint shapes any voice-agent budget.
Sub-500ms round-trip is the aggressive target, and most production use cases stay usable up to a higher practical ceiling. Hitting either requires component picks that compose to the budget.

The breakdown:
| Component | Typical range | Aggressive (sub-500ms) pick | Practical pick |
|---|---|---|---|
| STT | ~0.4s to first final | AssemblyAI Universal-3.5 Pro Realtime | Deepgram Nova-3 plus Flux |
| LLM inference | 100-200ms | Fast model (Gemini 3.7 Flash, Grok 4.6) | GPT-5.6 Sol, Claude models |
| TTS first audio | tens of ms vendor / higher p50 | Deepgram Flux TTS or Cartesia Sonic-3.6 | ElevenLabs Flash v2.5 or Aura-2 |
| Orchestration | 50-100ms | Platform-native | Standard platform |
Every time-to-first-audio in that table is vendor model latency measured in isolation. Independent p50 streaming latency runs higher, and a real stack adds network transit, so a live deployment sits nearer the top of each range.
That gap is exactly why this guide does not print a single round-trip number: any composite of best-case component figures would misrepresent what your users hear. For the LLM slot, a fast model keeps the loop tight, and the current low-latency picks are in the August LLM guide.
Cost at scale: what 100K minutes per month actually costs
Per-minute list price hides the real production cost, and the paths split into managed flat-fee, bring-your-own with passthrough, native token-metered audio, and self-hosted with engineering load. These figures carry from the July baseline, so re-verify before you commit budget:
| Stack | Pricing basis | Est. monthly (100K min) |
|---|---|---|
| Retell (managed) | from $0.07+/min, HIPAA on Enterprise | ~$7,000+ all-in |
| Vapi (BYO + passthrough) | $0.05/min fee; all-in ~$0.23-0.33/min | ~$23,000-33,000 |
| Deepgram-native (Nova-3 + Flux + TTS) | $0.0048/min STT + $0.0065/min Flux + $0.045/1k chars TTS | component sum plus LLM |
| Self-host (LiveKit / Pipecat + providers) | Permissive framework free plus provider passthrough | provider sum plus engineering |
| OpenAI Realtime (gpt-realtime-2.1) | per audio token | premium; scales with audio tokens |
The honest framing holds from prior months: under roughly 100K minutes per month, a managed platform such as Retell usually wins on total cost of ownership because the engineering time saved on passthrough tuning, retry logic, and compliance paperwork dominates the per-minute delta.
Above about 1M minutes per month, bring-your-own or self-hosted economics start to flip, and the crossover depends on your retry rate and engineering cost. Deepgram-native gets a boost through September 12 while Flux TTS is free.
Decision framework
Choose Cartesia Sonic-3.6 if:
- You want the highest independent TTS quality and sub-90ms vendor latency in one model.
- You need broad language coverage (44 languages).
- You can work with a beta and confirm final pricing.
Choose Deepgram Flux TTS (and Flux STT) if:
- You want a matched conversation-native STT and TTS pair from one vendor.
- Model-integrated turn-taking matters more than a long track record.
- You can adopt it before the free window closes on September 12.
Choose Microsoft MAI-Transcribe-1.5 if:
- You are on Azure or Foundry and want the best accuracy per unit of speed.
- Batch and near-real-time transcription at scale is the workload.
- You do not need model-native streaming turn-taking today.
Choose AssemblyAI Universal-3.5 Pro Realtime if:
- Your agent needs turn-by-turn context without reconnecting.
- You want structure on the streaming transcript, not just words.
- You can absorb $0.45 per hour for the accuracy.
Choose LiveKit or Pipecat over Vapi or Retell if:
- You want permissive open-source orchestration with no platform fee.
- You are willing to own reliability, retries, and on-call.
- You need first-party adapters across many STT and TTS providers.
Common mistakes when picking voice AI components in August 2026
-
Trusting vendor WER over an independent board. Self-reported accuracy runs optimistic. Read the neutral Artificial Analysis index, and remember batch WER does not predict streaming behavior.
-
Treating the LLM as free latency. A 0.4s STT paired with a 1,200ms LLM is not a fast agent. Pick a fast model for sub-500ms targets and measure the whole loop.
-
Reading vendor time-to-first-audio as end-to-end. A sub-100ms vendor figure is model latency in isolation. Real p50 adds network transit and orchestration, so budget from your own measurement.
-
Assuming HIPAA is default. Retell bundles a BAA on its Enterprise tier, and Vapi and LiveKit gate one behind an add-on or a plan. Confirm the compliance line before you build, not after.
-
Budgeting on a beta price or a free tier that expires. Cartesia Sonic-3.6 is in beta with pricing still settling, and Deepgram Flux TTS is only free through September 12. Price the standard tier before you commit.
How Future AGI fits
Picking STT and TTS on vendor word error rate and time-to-first-audio is not the same as knowing your call quality. Voice agents fail in production for the same reasons text agents do: the transcript drifts on an accent, the LLM hallucinates a policy, the agent talks over a caller, or a tool call fires on the wrong turn.
None of that shows up on a leaderboard, so the last step of any voice build is your own measurement loop.
Future AGI is the layer voice teams wrap around their framework of choice. Simulate generates voice conversations across the stack (accents, background noise, interruptions, ambiguous phrasing) and replays them against your agent before you ship.
You then evaluate the transcripts and outcomes with custom evals: write your own grading rule, choose an LLM judge or a deterministic check, set a pass or fail threshold, and run it as a gate in CI.
For voice specifically, that means Conversation Resolution and Task Completion to check the call actually achieved its goal, plus Tone and Groundedness on what the agent said, documented in the evaluation docs.
In production, observability traces live calls with OpenTelemetry-native instrumentation, so a failing turn surfaces as a scored span rather than a support ticket.
The core is Apache-2.0 and self-hostable via Docker Compose at no license cost, and the repo is the place to start. Future AGI is a companion to Vapi, Retell, LiveKit, and Pipecat, not a competitor on the voice-framework axis.
August 2026 settled into one clear truth for voice builders: sub-100ms streaming TTS is now table stakes, and every layer of the stack has at least two production-grade options. The stack you assemble matters less than the latency budget it lands inside and the eval loop you wrap around it.
Shortlist by your binding constraint, wire the STT, LLM, TTS, and orchestration path, then measure real p50 and p95 on your own accents and noise before you ship. That measurement loop, not the leaderboard row, is what turns August’s mature components into an agent that holds up on a live call.
Sources
STT primary
- AssemblyAI pricing (Universal-3.5 Pro, streaming and batch)
- AssemblyAI benchmarks
- Deepgram pricing (Nova-3, Flux STT)
- ElevenLabs Scribe v2 Realtime changelog (Aug 3)
- Microsoft MAI-Transcribe-1.5
TTS primary
- Deepgram Flux TTS (conversation-native, “as low as 80ms”)
- Cartesia Sonic-3.6 (beta, 44 languages, sub-90ms)
- ElevenLabs models (Flash v2.5, Eleven v3)
- Microsoft MAI-Voice-2
Voice agent platforms
Independent benchmarks and standards
- Artificial Analysis Text-to-Speech Arena (quality Elo)
- Artificial Analysis Speech-to-Text (independent WER)
- ITU-T Recommendation G.114 (one-way transmission time)
See also: Best LLMs of August 2026 for the LLM brain in your voice agent. Previous voice post: Best Voice AI of July 2026.
Frequently Asked Questions
What is the best text-to-speech model in August 2026?
What is the best speech-to-text model for voice agents in August 2026?
What changed in voice AI in August 2026?
What is the best voice agent platform in August 2026?
How much latency can a production voice agent tolerate in August 2026?
Which voice AI models are new in August 2026?
STT, TTS, and voice-agent picks for July 2026: ElevenLabs Scribe v2 leads accuracy at 2.2% WER, Deepgram Nova-3 for streaming, Cartesia Sonic-3.5 for TTS.
Best TTS APIs in May 2026: Cartesia Sonic 4 at 40ms, ElevenLabs v3, Deepgram Aura-2, Hume Octave, plus pricing, latency, and the right pick by use case.
Best STT APIs in May 2026: Deepgram Nova-3 + Flux, AssemblyAI Universal-2, Whisper, ElevenLabs Scribe v2 with WER, latency, and pricing compared.