Research

Best Voice AI Models in August 2026: STT, TTS, and Voice Agent Stack

August 2026 voice AI picks: Cartesia Sonic-3.6 tops both speech arenas, Deepgram Flux TTS ships sub-100ms, free through September, plus the STT and agent stack.

· 22 min read
voice-ai stt tts voice-agents monthly-compare 2026
Three-column map of the best voice AI stack for August 2026 on black: STT led by AssemblyAI Universal-3.5 Pro, TTS led by Cartesia Sonic-3.6, and voice agents led by Retell AI.
Table of Contents

You are shipping a voice agent, and the ground moved in August. Two conversation-native streaming text-to-speech models landed five days apart, both answering in under 100ms on vendor timing, and one of them took the top spot on both independent speech arenas within a week.

Streaming transcription runs turn-by-turn under half a second, a fast LLM closes the loop, and the hard part is choosing among components that all work. This guide picks the pieces of a production voice agent and the latency budget they must fit.

Best voice AI stack for August 2026 as a three-column map on black: an STT column led by AssemblyAI Universal-3.5 Pro over Deepgram Nova-3, Deepgram Flux, and ElevenLabs Scribe v2; a TTS column led by Cartesia Sonic-3.6 over Deepgram Flux TTS, ElevenLabs Flash v2.5, and Deepgram Aura-2; and a voice agents column led by Retell AI over Vapi, LiveKit Agents, Pipecat, and OpenAI Realtime, with an audio waveform strip across the top.

TL;DR: Best voice AI per layer, August 2026

LayerBest pickWhyPricing
TTS quality (independent)Cartesia Sonic-3.6#1 on both Artificial Analysis speech arenas (Aug 19), sub-90ms vendor TTFA~$49/1M chars; agents $0.06/min
TTS conversation-nativeDeepgram Flux TTSVendor “as low as 80ms”, model-integrated turns; free through Sept 12$0.045/1k chars
TTS expressiveEleven v370+ languages, audio tags, not real-time$0.10/1k chars
TTS low-latency (ElevenLabs)ElevenLabs Flash v2.5~75ms vendor inference, Turbo now deprecated, 40k-char limitPAYG char
TTS enterpriseDeepgram Aura-2Sub-200ms vendor steady-state$0.030/1k chars
STT streaming (agents)AssemblyAI Universal-3.5 Pro Realtime4.1% WER on AA-WER Streaming, ~0.4s to first final$0.45/hr ($7.50/1k min)
STT accuracy (batch)AssemblyAI Universal-3.5 Pro7.0% aggregate WER, promptable transcription$0.21/hr batch
STT accuracy-speedMicrosoft MAI-Transcribe-1.52.4% WER, #3 independent, leads the accuracy-speed frontier$6/1k min (Foundry)
STT cost / broad languagesDeepgram Nova-3Added Afrikaans and Georgian in August$0.0048/min
STT turn-takingDeepgram Flux (STT)Model-integrated end-of-turn under 400ms$0.0065/min EN
STT (ElevenLabs stacks)ElevenLabs Scribe v2 RealtimeAug 3: entity detection plus turbo and lite tiersPAYG
Voice agent (managed)Retell AISingle-number billing, HIPAA on the Enterprise tierfrom $0.07+/min
Voice agent at scale (BYO)Vapi$0.05/min platform fee, bring your own stack~$0.23-0.33/min all-in
Open-source orchestrationLiveKit Agents / PipecatApache-2.0 / BSD-2-Clause, self-host freeinfra + providers
Speech-to-speechOpenAI Realtime (gpt-realtime-2.1)Single-model audio loop (July release)per audio token
LLM brainSee the August LLM guidePick a fast model to keep the loop tightcross-reference

If you only read one row: Deepgram Nova-3 plus Flux for STT, a fast LLM brain such as Gemini 3.7 Flash or Grok 4.6 (see the August LLM guide), Cartesia Sonic-3.6 or Deepgram Flux TTS for speech, and Retell or Vapi to orchestrate.

That stack lands inside the practical round-trip budget and runs at production scale today.

Voice agent signal chain for August 2026 drawn left to right on black: a microphone feeds a streaming STT stage with AssemblyAI Universal-3.5 Pro, then an LLM brain stage with a fast model, then a TTS stage with Cartesia Sonic-3.6 highlighted bright white, then an orchestration stage with LiveKit and Pipecat, then a voice-reply output to a speaker, all under a bracket labeled ITU-T G.114 latency budget.

The story of voice AI in August 2026

Voice AI spent early 2026 turning research demos into datasheet-grade components, and August was the month text-to-speech caught up to the rest of the stack on latency. Two conversation-native streaming models launched inside a five-day window, both quoting sub-100ms time-to-first-audio, and the independent leaderboard confirmed one of them within days.

The engineering question is no longer whether a layer works. It is which pieces compose for your accents, languages, and cost ceiling.

The first launch was Deepgram Flux TTS on August 12, a conversation-aware model that quotes latency “as low as 80ms” even under load and handles turn-taking inside the model rather than leaving it to the application.

It lists at $0.045 per 1,000 characters after trial and is free through September 12, 2026, with 45 concurrent streams globally and 5 in the EU and Australia, moving to standard pricing on September 13.

Flux TTS pairs with Deepgram’s Flux STT, so a team can run a matched conversational STT and TTS pair from one vendor.

Five days later, Cartesia Sonic-3.6 shipped in beta on August 17, adding Odia and Urdu to reach 44 languages at a vendor time-to-first-audio under 90ms.

On August 18 to 19 it ranked first on both Artificial Analysis speech arenas, leading Provider Voice at roughly 1,283 Elo and Controlled Voice at roughly 1,123 Elo.

A vendor launch paired with an independent benchmark inside the same week is the cleanest signal the category produced in August, and it makes sub-100ms streaming TTS table stakes rather than a headline feature.

The beta pricing and exact general-availability date are still settling, so confirm the final list price on Cartesia’s own page.

The dated August news on transcription was ElevenLabs, whose Scribe v2 Realtime update on August 3 added entity detection, new scribe_v2_realtime_turbo and _lite IDs, and a secondary_languages parameter. Deepgram Nova-3 added Afrikaans and Georgian.

Microsoft MAI-Transcribe-1.5 still leads the accuracy-speed frontier at 2.4% word error rate, third on the raw independent board and up to five times faster than its peers, but it launched at Build on June 2, so it belongs on any shortlist as a standing pick rather than as August news.

The business story stayed quiet. No August funding, acquisition, or pricing move surfaced for Vapi, Retell, LiveKit, or Pipecat, and the reported ElevenLabs tender offer near a $22 billion valuation remains pending per July reporting, with its Series D closed back on February 4, 2026 at $11 billion led by Sequoia.

The round-trip target still anchors on ITU-T G.114, which sets one-way mouth-to-ear delay at 150ms preferred and 400ms tolerable. Hitting that budget in August 2026 is a matter of picking components that compose, then measuring your real p50 and p95 on your own traffic.

Best speech-to-text (STT) models in August 2026

STT split cleanly this month by job. One model leads raw accuracy per unit of speed, one leads streaming for agents, one leads cost and coverage, and one solves turn-taking. Batch word error rate does not predict streaming behavior, so match the benchmark to how you will actually run the model.

Microsoft MAI-Transcribe-1.5. The accuracy-speed leader

The first-party cloud STT that leads the accuracy-speed frontier, and a standing pick rather than an August launch, since it shipped at Build on June 2. MAI-Transcribe-1.5 posts 2.4% word error rate on the independent board, third overall on raw accuracy, while running up to five times faster than Gemini 3.1 Flash, Scribe v2, and gpt-4o-transcribe.

Specs:

  • WER: 2.4% on Artificial Analysis (independent); #3 on raw WER, first on the accuracy-speed Pareto frontier
  • Throughput: up to 5x faster than comparable models
  • Availability: Microsoft Foundry (launched June 2 at Build, not an August event)
  • Pricing: $6 per 1,000 minutes via Microsoft Foundry

Best for: Teams on Azure or Foundry that want the best accuracy per dollar per unit of speed, batch and near-real-time transcription at scale.

Skip if: You are not on Azure or Foundry, or you need model-native streaming turn-taking today (pair Nova-3 with Flux).

AssemblyAI Universal-3.5 Pro Realtime. Streaming with context

The pick when an agent needs turn-by-turn transcription without dropping context or reconnecting. Universal-3.5 Pro Realtime measures 4.1% word error rate on the AA-WER Streaming benchmark and returns a first final in roughly 0.4 seconds, and the batch sibling remains a high-accuracy async option.

Specs:

  • WER: 4.1% on AA-WER Streaming (independent) for the Realtime model; batch Universal-3.5 Pro at 7.0% aggregate WER
  • Latency: ~0.4s to first final, per-second billing
  • Pricing: $0.45/hr streaming ($7.50 per 1,000 minutes); $0.21/hr batch
  • Promptable transcription and structure on the streaming transcript

Best for: Multi-turn agents that need context carried across the call, multi-speaker calls, and domain jargon generic models mistranscribe.

Skip if: You need sub-$0.30 per hour transcription (use Deepgram Nova-3) or pure real-time dictation.

Deepgram Nova-3 and Flux STT. Cost, coverage, and turn-taking

The cost and coverage default, now with wider language support. Nova-3 added Afrikaans and Georgian in August and stays the streaming price leader, and Deepgram Flux STT adds model-integrated end-of-turn detection for agents that need to know when a speaker has actually finished.

Specs:

  • Nova-3: broad language coverage plus August additions (Afrikaans, Georgian), $0.0048/min streaming
  • Flux STT: model-integrated end-of-turn under 400ms, $0.0065/min EN
  • Both tuned for the sub-300ms budgets voice agents live in
  • Pairs natively with Deepgram Flux TTS for a matched conversational stack

Best for: Production agents where cost and language breadth are the binding constraints, and any turn-based agent that needs semantic end-of-turn rather than a fixed pause timer.

Skip if: You want the lowest independently measured accuracy figure (use MAI-Transcribe-1.5 or AssemblyAI) or pure dictation with no turn-taking.

ElevenLabs Scribe v2 Realtime. The August STT update

The dated August STT news for ElevenLabs stacks. The August 3 changelog added entity detection to Scribe v2 Realtime, new turbo and lite model IDs, and a secondary-languages parameter, keeping ElevenLabs transcription inside a live-agent latency window.

Specs:

  • August 3 update: entity detection, scribe_v2_realtime_turbo and _lite IDs, secondary_languages
  • Latency: realtime, tuned for live agents
  • Best fit inside an existing ElevenLabs voice stack
  • Pricing: pay as you go

Best for: Teams already standardized on ElevenLabs voices that want realtime transcription with structured entity output in the same stack.

Skip if: You need a published, independently measured word error rate to budget against (use MAI-Transcribe-1.5 or AssemblyAI).

Best text-to-speech (TTS) models in August 2026

August was the biggest TTS month of the year. Two conversation-native streaming models launched inside five days, both quoting sub-100ms time-to-first-audio, and one topped the independent quality board. Our text-to-speech providers guide tracks the broader field these two now lead.

One caution up front: every vendor time-to-first-audio below is model latency measured in isolation, not end-to-end, and independent p50 streaming latency runs higher.

Text-to-speech time-to-first-audio for August 2026 as horizontal bars where lower is better, marked vendor model latency not end-to-end: ElevenLabs Flash v2.5 at 75ms is the shortest bar and highlighted bright white, above Deepgram Flux TTS at 80ms, Cartesia Sonic-3.6 at 90ms, and Deepgram Aura-2 at 200ms.

Cartesia Sonic-3.6. The quality and latency leader

The structural pick when both naturalness and round-trip latency matter. Sonic-3.6 shipped in beta on August 17 with a vendor time-to-first-audio under 90ms, then ranked first on both Artificial Analysis speech arenas on August 18 to 19, so it leads quality and latency at the same time.

Specs:

  • Quality: #1 on Provider Voice (~1,283 Elo) and Controlled Voice (~1,123 Elo) on Artificial Analysis (Aug 18 to 19)
  • TTFA: sub-90ms vendor model latency, not end-to-end
  • Languages: 44 (added Odia and Urdu)
  • Pricing: ~$49 per 1M characters (aggregator-normalized), Scale $299/mo, voice agents $0.06/min; beta, confirm final pricing

Best for: The lowest-latency, most natural streaming TTS, telephony where latency drives perceived call quality, real-time interactive apps.

Skip if: You need open weights (none) or a settled general-availability price today (it is still in beta).

Deepgram Flux TTS. Conversation-native TTS

The conversation-aware TTS launched August 12, built to handle turn-taking inside the model rather than at the application layer. Flux TTS quotes latency “as low as 80ms” even under load, and it pairs with Deepgram Flux STT for a matched conversational STT and TTS pair from one vendor.

Specs:

  • Latency: “as low as 80ms” vendor model latency, even under load
  • Turn handling: model-integrated, conversation-aware
  • Pricing: $0.045 per 1,000 characters after trial; free through September 12, 2026 (45 concurrent streams global, 5 EU and Australia); standard from September 13
  • Pairs with Deepgram Flux STT for a matched stack

Best for: Enterprise agents that want a single-vendor conversational STT and TTS pair, and teams that can adopt it before the free window closes.

Skip if: You need it free past September 12, or you want a long production track record (it is brand-new).

ElevenLabs Flash v2.5 and Eleven v3. The ElevenLabs paths

The two ElevenLabs lanes. Flash v2.5 is the real-time conversational path near 75ms model inference, and in August ElevenLabs deprecated its Turbo models and named Flash the replacement, with a 40,000-character request limit. Eleven v3 remains the expressive model for non-real-time work.

Specs:

  • Flash v2.5: ~75ms vendor model inference, Turbo deprecated in August, 40,000-character request limit
  • Eleven v3: 70+ languages, audio tags, expressive; generally available since February 2026, not real-time
  • Pricing: Flash v2.5 pay as you go, character-based; Eleven v3 at $0.10 per 1,000 characters
  • Voice cloning supported

Best for: Real-time conversational agents that want ElevenLabs voices (Flash v2.5); audiobooks, narration, and character work rendered ahead of time (Eleven v3).

Skip if: You need the lowest possible time-to-first-audio (use Cartesia Sonic-3.6) or real-time output from v3 (use Flash v2.5).

Deepgram Aura-2. Enterprise TTS

The pick when an existing Deepgram enterprise deployment wants first-party TTS end to end. Aura-2 carries a broad voice set at a vendor sub-200ms steady-state latency, which fits teams already running Nova-3 and Flux STT.

Specs:

  • Latency: sub-200ms vendor steady-state; higher on independent p50 streaming
  • Pricing: $0.030 per 1,000 characters (pay as you go)
  • First-party fit with the Deepgram STT stack

Best for: Existing Deepgram enterprise users that want one vendor for STT and TTS.

Skip if: You want the newer conversation-native Flux TTS, or the highest independent quality score (use Sonic-3.6).

Microsoft MAI-Voice-2. The hyperscaler expressive path

The Foundry expressive TTS line, and like MAI-Transcribe a standing pick rather than an August launch, since it reached general availability at Build on June 2. MAI-Voice-2 covers 15 languages with short-sample cloning, and the MAI-Voice-2-Flash preview is the faster, lower-cost variant.

Specs:

  • MAI-Voice-2: expressive TTS, 15 languages, short-sample voice cloning (GA June 2 at Build)
  • MAI-Voice-2-Flash: preview, “2x speed at 32% lower cost” per vendor
  • Availability: Microsoft Foundry
  • Pricing: Foundry pricing, not fully re-confirmed

Best for: Azure, Foundry, and Copilot stacks that want first-party expressive TTS.

Skip if: You are not on Azure, or you need a published, independently measured time-to-first-audio to budget against (use Cartesia or ElevenLabs).

Best voice agent platforms in August 2026

If you do not want to wire STT, LLM, TTS, and orchestration yourself, the platforms below ship in days what a custom build ships in quarters. No August pricing, funding, or compliance change surfaced for any of them, so treat the numbers here as the July baseline and confirm the current tier before you build.

Retell AI. The most-teams default

The right default for most production voice-agent teams. Retell bundles HIPAA and a BAA on its Enterprise tier and bills through a single number, which removes the passthrough and compliance overhead that slows a build.

Specs:

  • Pricing: from $0.07+ per minute, single-vendor billing
  • Builder: no-code visual builder plus developer SDK
  • Compliance: HIPAA and BAA bundled (Enterprise tier; re-verify)

Best for: Most production voice agents that want managed HIPAA and a platform that reduces engineering load.

Skip if: You want to bring your own every component (use Vapi) or a fully self-hosted stack (use LiveKit or Pipecat).

Vapi. The scale and BYO pick

The pick when you want to bring your own STT, LLM, and TTS and run at volume. Vapi charges a $0.05 per minute platform fee on top of component passthrough, landing all-in around $0.23 to $0.33 per minute.

Specs:

  • Pricing: $0.05/min platform fee plus passthrough; all-in ~$0.23-0.33/min BYOK
  • Bring-your-own STT, LLM, and TTS
  • Compliance: HIPAA BAA add-on at $2,000/mo (baseline; re-verify)

Best for: Teams that want component-level control and volume scale.

Skip if: You are cost-sensitive at scale or need a cheap BAA (Retell bundles it), or you want a self-hosted stack.

OpenAI Realtime API. Native speech-to-speech

The pick when you want one model for the whole audio loop. The Realtime API runs gpt-realtime-2.1, which released July 6 and cut p95 latency by at least 25% versus the prior model, priced per audio token rather than per minute.

Specs:

  • Model: gpt-realtime-2.1 and mini, speech-to-speech (July 6 release, current in August)
  • Latency: p95 down at least 25% versus prior
  • Pricing: per audio token; mini priced as the prior gpt-realtime-mini

Best for: Prosody-sensitive products and single-vendor OpenAI stacks that value a simpler loop over per-minute predictability.

Skip if: You need per-minute cost control, or separable STT and TTS control.

LiveKit Agents and Pipecat. Open-source orchestration

The open-source route when you want to own the orchestration and still get first-party adapters for every STT and TTS. LiveKit Agents is Apache-2.0 with a managed Cloud, and Pipecat is BSD-2-Clause and fully self-composed.

Specs:

  • LiveKit Agents: Apache-2.0, self-host free; Cloud adds $0.01/min agent sessions; signed BAA at Scale ($500/mo, baseline)
  • Pipecat: BSD-2-Clause framework, free; plugin ecosystem across Cartesia, Deepgram, ElevenLabs, and OpenAI
  • Compliance: LiveKit BAA at Scale; Pipecat is your own deployment boundary

Best for: Teams that want permissive open-source orchestration with production infrastructure (LiveKit) or full control of the orchestration code (Pipecat).

Skip if: You want a fully managed no-code platform (use Retell) or a signed BAA below the Scale tier.

HIPAA and compliance tier

Compliance is the cleanest differentiator across managed platforms, and every signed BAA sits behind an enterprise plan or a paid add-on. These figures carry from July, so confirm the exact tier before you build:

PlatformHIPAA / signed BAAWhere it lands
Retell AIBundledHIPAA and BAA included in the flat rate (Enterprise; re-verify)
Vapi$2,000/mo add-onPlatform fee plus passthrough on top
LiveKit AgentsScale tierSigned BAA at Scale ($500/mo baseline)
Pipecat / self-hostYour own deploymentYou own the compliance boundary

Vendor WER versus independent benchmarks

Every STT vendor publishes a best-case word error rate, and those numbers run optimistic against a neutral board. The independent Artificial Analysis index is the cleaner cross-provider read, our speech-to-text APIs guide breaks the providers down on pricing and benchmarks, and batch accuracy does not predict streaming behavior.

On the TTS side, vendor time-to-first-audio is a floor, not what your users hear, so a domain reproduction with your accents, your background noise, and your prompts is the only number that decides the pick.

End-to-end latency budget. The math

Voice agents have a hard latency target. The ITU-T G.114 one-way mouth-to-ear recommendation is 150ms preferred and 400ms tolerable, and that constraint shapes any voice-agent budget.

Sub-500ms round-trip is the aggressive target, and most production use cases stay usable up to a higher practical ceiling. Hitting either requires component picks that compose to the budget.

Voice agent end-to-end latency budget for August 2026 as a single stacked bar of vendor best-case component latencies: streaming STT around 400ms with AssemblyAI Universal-3.5 Pro, LLM inference around 150ms with a fast model, TTS around 80 to 90ms with Deepgram Flux TTS or Cartesia Sonic-3.6, and orchestration around 50ms, measured against a sub-700ms practical target and well above the 150ms ITU-T G.114 one-way ideal, with real p50 latency higher.

The breakdown:

ComponentTypical rangeAggressive (sub-500ms) pickPractical pick
STT~0.4s to first finalAssemblyAI Universal-3.5 Pro RealtimeDeepgram Nova-3 plus Flux
LLM inference100-200msFast model (Gemini 3.7 Flash, Grok 4.6)GPT-5.6 Sol, Claude models
TTS first audiotens of ms vendor / higher p50Deepgram Flux TTS or Cartesia Sonic-3.6ElevenLabs Flash v2.5 or Aura-2
Orchestration50-100msPlatform-nativeStandard platform

Every time-to-first-audio in that table is vendor model latency measured in isolation. Independent p50 streaming latency runs higher, and a real stack adds network transit, so a live deployment sits nearer the top of each range.

That gap is exactly why this guide does not print a single round-trip number: any composite of best-case component figures would misrepresent what your users hear. For the LLM slot, a fast model keeps the loop tight, and the current low-latency picks are in the August LLM guide.

Cost at scale: what 100K minutes per month actually costs

Per-minute list price hides the real production cost, and the paths split into managed flat-fee, bring-your-own with passthrough, native token-metered audio, and self-hosted with engineering load. These figures carry from the July baseline, so re-verify before you commit budget:

StackPricing basisEst. monthly (100K min)
Retell (managed)from $0.07+/min, HIPAA on Enterprise~$7,000+ all-in
Vapi (BYO + passthrough)$0.05/min fee; all-in ~$0.23-0.33/min~$23,000-33,000
Deepgram-native (Nova-3 + Flux + TTS)$0.0048/min STT + $0.0065/min Flux + $0.045/1k chars TTScomponent sum plus LLM
Self-host (LiveKit / Pipecat + providers)Permissive framework free plus provider passthroughprovider sum plus engineering
OpenAI Realtime (gpt-realtime-2.1)per audio tokenpremium; scales with audio tokens

The honest framing holds from prior months: under roughly 100K minutes per month, a managed platform such as Retell usually wins on total cost of ownership because the engineering time saved on passthrough tuning, retry logic, and compliance paperwork dominates the per-minute delta.

Above about 1M minutes per month, bring-your-own or self-hosted economics start to flip, and the crossover depends on your retry rate and engineering cost. Deepgram-native gets a boost through September 12 while Flux TTS is free.

Decision framework

Choose Cartesia Sonic-3.6 if:

  • You want the highest independent TTS quality and sub-90ms vendor latency in one model.
  • You need broad language coverage (44 languages).
  • You can work with a beta and confirm final pricing.

Choose Deepgram Flux TTS (and Flux STT) if:

  • You want a matched conversation-native STT and TTS pair from one vendor.
  • Model-integrated turn-taking matters more than a long track record.
  • You can adopt it before the free window closes on September 12.

Choose Microsoft MAI-Transcribe-1.5 if:

  • You are on Azure or Foundry and want the best accuracy per unit of speed.
  • Batch and near-real-time transcription at scale is the workload.
  • You do not need model-native streaming turn-taking today.

Choose AssemblyAI Universal-3.5 Pro Realtime if:

  • Your agent needs turn-by-turn context without reconnecting.
  • You want structure on the streaming transcript, not just words.
  • You can absorb $0.45 per hour for the accuracy.

Choose LiveKit or Pipecat over Vapi or Retell if:

  • You want permissive open-source orchestration with no platform fee.
  • You are willing to own reliability, retries, and on-call.
  • You need first-party adapters across many STT and TTS providers.

Common mistakes when picking voice AI components in August 2026

  1. Trusting vendor WER over an independent board. Self-reported accuracy runs optimistic. Read the neutral Artificial Analysis index, and remember batch WER does not predict streaming behavior.

  2. Treating the LLM as free latency. A 0.4s STT paired with a 1,200ms LLM is not a fast agent. Pick a fast model for sub-500ms targets and measure the whole loop.

  3. Reading vendor time-to-first-audio as end-to-end. A sub-100ms vendor figure is model latency in isolation. Real p50 adds network transit and orchestration, so budget from your own measurement.

  4. Assuming HIPAA is default. Retell bundles a BAA on its Enterprise tier, and Vapi and LiveKit gate one behind an add-on or a plan. Confirm the compliance line before you build, not after.

  5. Budgeting on a beta price or a free tier that expires. Cartesia Sonic-3.6 is in beta with pricing still settling, and Deepgram Flux TTS is only free through September 12. Price the standard tier before you commit.

How Future AGI fits

Picking STT and TTS on vendor word error rate and time-to-first-audio is not the same as knowing your call quality. Voice agents fail in production for the same reasons text agents do: the transcript drifts on an accent, the LLM hallucinates a policy, the agent talks over a caller, or a tool call fires on the wrong turn.

None of that shows up on a leaderboard, so the last step of any voice build is your own measurement loop.

Future AGI is the layer voice teams wrap around their framework of choice. Simulate generates voice conversations across the stack (accents, background noise, interruptions, ambiguous phrasing) and replays them against your agent before you ship.

You then evaluate the transcripts and outcomes with custom evals: write your own grading rule, choose an LLM judge or a deterministic check, set a pass or fail threshold, and run it as a gate in CI.

For voice specifically, that means Conversation Resolution and Task Completion to check the call actually achieved its goal, plus Tone and Groundedness on what the agent said, documented in the evaluation docs.

In production, observability traces live calls with OpenTelemetry-native instrumentation, so a failing turn surfaces as a scored span rather than a support ticket.

The core is Apache-2.0 and self-hostable via Docker Compose at no license cost, and the repo is the place to start. Future AGI is a companion to Vapi, Retell, LiveKit, and Pipecat, not a competitor on the voice-framework axis.

August 2026 settled into one clear truth for voice builders: sub-100ms streaming TTS is now table stakes, and every layer of the stack has at least two production-grade options. The stack you assemble matters less than the latency budget it lands inside and the eval loop you wrap around it.

Shortlist by your binding constraint, wire the STT, LLM, TTS, and orchestration path, then measure real p50 and p95 on your own accents and noise before you ship. That measurement loop, not the leaderboard row, is what turns August’s mature components into an agent that holds up on a live call.

Sources

STT primary

TTS primary

Voice agent platforms

Independent benchmarks and standards


See also: Best LLMs of August 2026 for the LLM brain in your voice agent. Previous voice post: Best Voice AI of July 2026.

Frequently Asked Questions

What is the best text-to-speech model in August 2026?

Cartesia Sonic-3.6 (beta, August 17) topped both Artificial Analysis speech arenas on August 18 to 19 at a vendor sub-90ms time-to-first-audio. Deepgram Flux TTS (August 12) is the conversation-native pick and is free through September 12, then moves to standard pricing on September 13.

What is the best speech-to-text model for voice agents in August 2026?

AssemblyAI Universal-3.5 Pro Realtime leads for turn-by-turn agents at 4.1% streaming word error rate and roughly 0.4s to first final. Deepgram Nova-3 wins on cost and language breadth, and Microsoft MAI-Transcribe-1.5 leads accuracy per unit of speed.

What changed in voice AI in August 2026?

The headline was speech synthesis. Two conversation-native, sub-100ms streaming TTS models launched five days apart: Deepgram Flux TTS on August 12 and Cartesia Sonic-3.6 on August 17. Sonic-3.6 then ranked first on both Artificial Analysis speech arenas on August 18 to 19.

What is the best voice agent platform in August 2026?

Retell AI and Vapi lead managed deployment, and LiveKit Agents (Apache-2.0) and Pipecat (BSD-2-Clause) lead open-source self-host. No August pricing or funding news surfaced, so treat BAA and per-minute figures as July baseline and re-verify the current tier before you commit budget.

How much latency can a production voice agent tolerate in August 2026?

ITU-T G.114 sets one-way mouth-to-ear delay at 150ms preferred and up to 400ms tolerable. Vendor time-to-first-audio is a best-case floor, not end-to-end, so measure your own p50 and p95 on real traffic, including network transit, before you commit to a stack.

Which voice AI models are new in August 2026?

Three launches shipped in August: Deepgram Flux TTS (August 12), Cartesia Sonic-3.6 beta (August 17), and ElevenLabs Scribe v2 Realtime tiers (August 3). Microsoft MAI-Transcribe-1.5 and MAI-Voice-2 are current top picks but launched at Build on June 2, not in August.
Related Articles
View all