Best LLMs of August 2026: Top Closed-Source, Open-Weight, Coding, and Multimodal Picks
Best LLMs of August 2026 by use case: Claude Opus 5 for agents, Grok 4.6 for cheap frontier reasoning, Gemini 3.7 Flash for coding, Muse Glimmer for local.
Table of Contents
Series note. This is the August 2026 entry in our monthly best-LLMs series. Google’s flagship stayed unshipped while the open-weight tier escalated and the frontier repriced again. Previous: July 2026. The closed frontier repriced and open weights hit a scale record. June 2026. The open-weight coding price war went mainstream.

TL;DR: Best LLM per category, August 2026
| Use case | Best pick | Why | Output $/M tokens |
|---|---|---|---|
| Highest intelligence (independent) | Claude Opus 5 (1M context) | Tops the Artificial Analysis Intelligence Index at 63.0 on the Aug 31 snapshot | $25 |
| Cheapest seat at the frontier | xAI Grok 4.6 | AA Index 60.9 at $2/$6 under 200K, the lowest price at the top | $6 |
| Coding value | Google Gemini 3.7 Flash | Shipped Aug 13, DeepSWE v1.1 65.3 (provider), intro rate through Dec 31 | $3.75 (intro) |
| Open-weight coding scale | Alibaba Qwen3.8-Max | SWE-bench Pro 67.7 (provider), but custom license, not Apache/MIT | ~$6 (aggregator) |
| Frontier reasoning after the cut | GPT-5.6 Sol | Price cut Aug 21 to $4/$20, top-handful on the index | $20 |
| Open-weight scale leader | Kimi K3 (Modified MIT) | Top open model on the AA Index at 59.7, 2.8T parameters | $15 |
| Best small / local open | Meta Muse Glimmer (Apache-2.0) | 30B, Meta’s first OSI-licensed model, runs on one GPU | free (open weights) |
| Cheapest permissive laptop model | Alibaba Qwen3.8-27B (Apache-2.0) | Consumer-hardware model, matches larger models (provider) | free (open weights) |
| Long-context agent coder | DeepSeek V4-Pro (MIT) | 1.6T MoE, 1M context, off-peak pricing effective Aug 16 | $1.98 (off-peak) |
| Cheap permissive frontier preview | Tencent Hunyuan Hy4 (Apache-2.0) | 770B MoE open-weight preview, priced below the closed frontier | $2.50 |
| Open coding / cyber | Z.ai GLM-5.3 | DeepSWE v1.1 66.9 (provider), flagship license unstated | open weights |
| Highest-capability Claude | Claude Fable 5 | AA Index 62.1, the top Claude tier | $50 |
| Restricted enterprise security | Claude Mythos 5 | Moved into Claude Security Aug 21 for enterprise defenders | not a standard API tier |
If you only read one row: Claude Opus 5 for the top independent score, Grok 4.6 for the cheapest frontier seat at $2/$6, Gemini 3.7 Flash for coding value, Meta Muse Glimmer for an Apache-2.0 model that runs on one GPU, and DeepSeek V4-Pro for 1M-context agent coding. Everything else trades around those five.

The single biggest story of August 2026: Google’s flagship did not ship. As of August 31, Gemini 3.5 Pro remained unreleased more than three months after its May announcement, and Google’s actual August release was the mid-tier Gemini 3.7 Flash on August 13.
SemiAnalysis reported that Google had shelved 3.5 Pro and pivoted to Gemini 4, but Google publicly disputes that account and says the model is in partner testing.
So the undisputed fact is narrow: the flagship is still not out, and a workhorse shipped in its place. Everything else in the market moved to fill that vacuum, from an open-weight surge to another round of price cuts at the top.
The story of August 2026: a missing flagship, an open-weight surge, and another price cut
July was the closed frontier repricing itself. August was the month a flagship failed to arrive and everyone else moved into the gap. Four things happened at once, and each one changes a production decision.
Google’s absence set the tone. Gemini 3.5 Pro was announced in May and, as of August 31, still had not shipped, so Google’s real August release was Gemini 3.7 Flash on August 13, a mid-tier coding and agent workhorse at an intro rate of $0.75/$3.75 per million tokens through the end of 2026.
The flagship the workhorse was meant to sit beneath never came.
The open-weight tier surged into the space. Meta shipped Muse Glimmer on August 10 under Apache-2.0, its first OSI-licensed open model, a 30B that runs on a single GPU.
Alibaba put Qwen3.8-Max weights up on August 12, Z.ai released GLM-5.3 (API August 14, weights August 28), and Tencent posted a Hunyuan Hy4 preview on August 28. Several of these weight drops ran two weeks behind their API launches, delayed over cyber-capability reviews.
The price war escalated again. On August 21 OpenAI cut GPT-5.6 Sol from $5/$30 to $4/$20, a promotional rate through about November 21. Grok 4.6 had already launched on August 12 at $2/$6 under 200K tokens, the cheapest seat at the frontier. The top of the market keeps getting cheaper faster than the bottom.

DeepSeek closed the loop on economics. V4-Pro reached formal general availability on August 12 to 13, and on August 16 DeepSeek introduced peak and off-peak API pricing, with off-peak running at roughly half the peak rate. For batch and asynchronous agent work that can be scheduled, the effective cost of a 1M-context MIT model dropped by half overnight.
The takeaway from August 2026: the flagship race stalled at the very top while the open and value tiers did all the moving. Model choice is still a price, license, and reliability decision, and the widening gap between a provider benchmark and your production number is what decides whether the cheap pick or the capable one actually holds.
Top closed-source / proprietary LLMs in August 2026
Claude Opus 5. Best for balanced agentic work
Anthropic. Generally available July 24, 2026. Carries into August as the independent-index leader.
Opus 5 tops the Artificial Analysis Intelligence Index at 63.0 on the August 31 snapshot, the highest independent score of any model this month. Anthropic had no new August model event, so its lineup held while rivals shipped. It stays the safe default for complex agentic coding and enterprise work.
- AA Intelligence Index 63.0, first overall (independent, Aug 31 snapshot)
- 1M-token context window (carried from July, verify live)
- $5 input / $25 output per million tokens, unchanged since July
What it does not win: price. At $25 output it is four times Grok 4.6 and nearly seven times Gemini 3.7 Flash, and two open-weight models sit within two points of it on the index.
OpenAI GPT-5.6 Sol. Best for frontier reasoning after the price cut
OpenAI. Base generally available July 9, 2026. Price cut August 21, 2026.
On August 21 OpenAI cut Sol from $5/$30 to $4/$20 per million tokens, a promotional rate that runs to about November 21. That moves a flagship reasoning tier into the same price band as much smaller models from a year ago. Sol still sits in the top handful on the independent index.
- Price cut to $4 input / $20 output per million tokens (cached $0.40), promo through about Nov 21
- AA Intelligence Index 58.9 on the Aug 31 snapshot, though some mirrors put it closer to 61
- Requests above roughly 272K tokens bill at 2x input and 1.5x output (verify live)
What it does not win: the top of the index. Claude Opus 5 and Fable 5 still outscore it, and Grok 4.6 matches its neighborhood for a third of the output price.
Google Gemini 3.7 Flash. Best for low-cost coding and agent workloads
Google DeepMind. Generally available August 13, 2026. API GA August 27.
This is the model Google shipped instead of Gemini 3.5 Pro. It is a mid-tier workhorse aimed at coding and agentic fan-out, and its introductory pricing is the cheapest at the coding frontier. The intro rate holds through the end of 2026 before it roughly doubles.
- $0.75 input / $3.75 output per million tokens (intro through Dec 31 2026), then $1.50/$7.50 from Jan 1 2027, both provider-confirmed
- DeepSWE v1.1 65.3, FrontierCode 1.1 43.6, WebDev Arena Elo 1588 (provider-reported)
- 1,048,576-token context per aggregators, though Google’s own pages omit the window
What it does not win: raw capability. It is the workhorse, not the flagship, and the flagship it was meant to sit beneath never shipped this month.
xAI Grok 4.6. Best for the cheapest seat at the frontier
xAI. Generally available August 12, 2026.
Grok 4.6 is the month’s clearest value story. It lands near the top of the independent index while pricing under every other frontier model. Under 200K tokens it costs $2/$6 per million, with cached input at $0.50, and it supersedes Grok 4.5.
- AA Intelligence Index 60.9 on the Aug 31 snapshot, third overall
- $2 input / $6 output under 200K tokens, $4/$12 above 200K, both provider release-notes confirmed
- CursorBench v3.2 69.9, DeepSWE v1.1 65.9 (provider-reported)
What it does not win: a published context number. The 500K window is aggregator-only and x.ai omits it, so confirm it before you design around long inputs.
Claude Mythos 5. Best for restricted enterprise security work
Anthropic. Brought into Claude Security August 21, 2026.
Mythos 5 shares Fable 5’s base and was invite-only until August 21, when Anthropic moved it into Claude Security for enterprise defenders alongside a $35M defender fund. It is a safety-gated model, not a standard API tier, and Anthropic frames it for defensive security work.
- Moved into Claude Security for enterprise on Aug 21
- Same base model as Claude Fable 5
- Safety-gated access rather than general API pricing
What it does not win: general availability. You cannot call it like a normal model, and Anthropic’s own materials describe defensive use, so do not assume offensive capability.
Top open-weight and Chinese-frontier LLMs in August 2026
If you want the best open source LLM this month, the answer splits by what you mean by open. One of these is the largest open model shipped, one is the first Apache-2.0 model from Meta, and two carry licenses that are not what “open weights” implies.
Meta Muse Glimmer. Best for local and small-footprint open models
Meta. Open weights August 10, 2026. Apache-2.0.
This is the real strategic shift of the month. Muse Glimmer is Meta’s first OSI-licensed open-weight model, a 30B distilled from Muse Spark that runs on a single consumer GPU. Apache-2.0 means commercial use without the Llama license carve-outs.
- Apache-2.0, about 29.6B parameters (marketed 30B), 131K+ context
- SWE-Bench Pro 51.2, AIME 2026 94.7, GAIA2 43.3 (provider-reported)
- Runs in roughly 17 to 20GB of VRAM at 4-bit
What it does not win: raw frontier scores. It is built to run cheaply and locally, not to top the index, and larger open models beat it on hard coding.
Alibaba Qwen3.8-Max. Best for open-weight coding scale
Alibaba. Open weights around August 12, 2026. Custom Qwen3.8-Max license.
Qwen3.8-Max is now downloadable, the first Max-class Qwen you can host yourself. It reports the strongest open coding numbers of the month, but read the license before you ship. It is a custom Qwen3.8-Max license with usage carve-outs, not Apache or MIT.
- 2.4T total / 95B active mixture-of-experts, 262K native context, extensible toward 1M
- SWE-bench Pro 67.7, Terminal Bench 2.1 86.6, GPQA Diamond 92.6 (provider-reported); AA Index 58.1
- Custom license, not Apache/MIT; the separate Qwen3.8-27B is Apache-2.0
What it does not win: a clean permissive license. If you need Apache or MIT terms, the small Qwen3.8-27B or Muse Glimmer fit, not this flagship.
Z.ai GLM-5.3. Best for open coding and cyber-reasoning
Z.ai. API August 14, 2026. Weights August 28. Flagship license unstated.
GLM-5.3 is the same 744B base as GLM-5.2 with post-training gains, and Z.ai calls it open-source coding state of the art. The weights shipped two weeks after the API, delayed over a cyber-capability review. One caution before you plan around it: Z.ai has not stated the flagship’s license.
- DeepSWE v1.1 66.9, Terminal-Bench 3.0 28.3, CyberGym 84.5 (provider-reported); AA Index 59.5
- 1,000,000-token context, 128K output
- GLM-5.2 and GLM-5.3-Flash shipped MIT, but the full GLM-5.3 flagship license is unstated
What it does not win: license certainty. Until Z.ai states the flagship terms, treat it as unresolved rather than assuming MIT.
DeepSeek V4-Pro. Best for long-context agent coding on a budget
DeepSeek. Formal general availability August 12 to 13, 2026. MIT.
V4-Pro reached formal GA in mid-August, and on August 16 DeepSeek introduced peak and off-peak API pricing. Off-peak runs at roughly half the peak rate, so batch and asynchronous agent work gets materially cheaper if you can time it. It is MIT-licensed with native OpenAI Responses API support.
- 1.6T total / 49B active mixture-of-experts, 1M context, up to 384K output
- Off-peak $0.66/$1.98, peak $1.32/$3.96 per million tokens, effective Aug 16 at 16:00 UTC
- Terminal-Bench 2.1 87.9, HLE 42.7 without tools (provider-reported)
What it does not win: peak-hour price predictability. The two-tier rate rewards scheduling, so ad-hoc daytime traffic pays the higher number.
Moonshot Kimi K3. Best for the largest open-weight model
Moonshot. Open weights July 27, 2026. Modified MIT. Carries into August.
Kimi K3 had no August event, but it stays the open-weight leader on the independent index and the largest open model shipped. At 2.8T parameters it is not cheap to host, so the download is free in license, not in inference cost.
- AA Intelligence Index 59.7, the top open-weight model on the board
- 2.8T parameters, about 1M context, FrontierSWE 81.2 (provider-reported)
- $3/$15 per million tokens on a hosted endpoint if you skip self-hosting
What it does not win: cheap local hosting. Its size makes it a data-center model, so Muse Glimmer or Qwen3.8-27B fit a single box better.
Tencent Hunyuan Hy4 (preview). Best for a cheap permissive frontier preview
Tencent. Open-weight preview August 28, 2026. Apache-2.0.
Hy4 arrived as an Apache-2.0 open-weight preview at the end of the month, a 770B mixture-of-experts model priced well below the closed frontier. It is a preview, not general availability, so treat the numbers as early and expect changes before a stable release.
- 770B total / 49B active mixture-of-experts, over 1M context
- $0.834 / $2.501 per million tokens (cached $0.042)
- Provider internal blind eval 2.99 out of 4, ahead of GLM-5.3 and Kimi K3 on that internal test
What it does not win: production readiness. Preview status means scores and endpoints can move, so pin a version and re-test before you rely on it.
Top multimodal LLMs in August 2026
August multimodal news was mostly platform housekeeping, not new leaderboards. Three changes on Google’s stack matter for anyone building on it.
| Date | Change | What it means |
|---|---|---|
| Aug 17 | Google Imagen 4 hard shutdown (imagen-4.0-generate-001, -ultra, -fast) | Image generation steers to gemini-3.1-flash-image (“Nano Banana”) |
| Aug 26 | Google Gemini 3.5 Transcribe GA | Speech-to-text, 85+ languages, diarization, word-level timestamps |
| Aug 27 | Google Gemini Omni Flash GA (gemini-omni-1.1-flash) | Video extension and first-plus-last-frame interpolation up to 4K |
On the open side, DeepSeek shipped V4-Flash-Vision-Exp on August 21 as an experimental multimodal variant, not a general-availability model, so keep it in testing.
The Imagen 4 shutdown is the one that breaks pipelines: the three endpoints are gone, and Google now routes image requests to its Gemini image model, so update any code still calling the old Imagen paths before it fails.
Voice and audio (covered separately)
Speech and voice models get their own guide. August was a big month for text-to-speech, with two sub-100ms streaming launches inside five days, so see Best Voice AI Models in August 2026 for the STT, TTS, and voice-agent picks.
Top embeddings and retrieval models in August 2026
No dated August 2026 embeddings release surfaced, so the picks carry from earlier in the year. This is a case where the honest answer is “no notable change,” and manufacturing one would only mislead.
| Model | Provider | Use case |
|---|---|---|
| Gemini Embedding 2 | General-purpose retrieval, strong multilingual | |
| voyage-context-4 / voyage-4-large | Voyage AI | Context-aware retrieval for RAG |
| gemini-embedding-001 | Stable production embedding endpoint | |
| text-embedding-3-large | OpenAI | OpenAI-native retrieval |
| embed-v4 | Cohere | Enterprise retrieval and reranking |
| Qwen3-Embedding-8B | Alibaba | Open-weight embedding for self-host |
Top coding-specific LLMs in August 2026
The best LLM for coding in August depends on whether you optimize for price, license, or raw score. Every number below is provider-reported, because independent SWE-bench Verified rankings were not confirmable this month and the public sources conflict.
| Model | Coding benchmark (provider-reported) | Context | Output $/M | License |
|---|---|---|---|---|
| Google Gemini 3.7 Flash | DeepSWE v1.1 65.3, FrontierCode 1.1 43.6 | 1M (aggregator) | $3.75 intro | Closed |
| xAI Grok 4.6 | CursorBench v3.2 69.9, DeepSWE v1.1 65.9 | 500K (aggregator) | $6 | Closed |
| Alibaba Qwen3.8-Max | SWE-bench Pro 67.7, Terminal Bench 2.1 86.6 | 262K | ~$6 (aggregator) | Custom |
| Z.ai GLM-5.3 | DeepSWE v1.1 66.9, Terminal-Bench 3.0 28.3 | 1M | open weights | Unstated |
| DeepSeek V4-Pro | Terminal-Bench 2.1 87.9 | 1M | $1.98 (off-peak) | MIT |
Do not read across these rows as if the benchmarks were the same test. They are not, and there is no single trustworthy “SWE-bench Verified #1” this month, so reproduce any score that matters on 100 to 500 of your own prompts before you budget around it. Our LLM evaluation tools guide covers how to build that check.
Top reasoning and math LLMs in August 2026
The independent Artificial Analysis Intelligence Index is the cleanest cross-provider read on reasoning. This is the August 31 snapshot, and it is a living leaderboard, so re-pull the exact values before you quote a decimal.
| Rank | Model | AA Intelligence Index (independent) |
|---|---|---|
| 1 | Claude Opus 5 | 63.0 |
| 2 | Claude Fable 5 | 62.1 |
| 3 | xAI Grok 4.6 | 60.9 |
| 4 | Moonshot Kimi K3 | 59.7 |
| 5 | Z.ai GLM-5.3 | 59.5 |
| 6 | OpenAI GPT-5.6 Sol | 58.9 |
| 7 | Qwen3.8-Max | 58.1 |
| 8 | GLM-5.3-Flash | 57.5 |
| 9 | Claude Opus 4.8 | 57.3 |
| 10 | Meta Muse Spark 1.2 | 56.8 |
Claude Opus 5 and Fable 5 hold the top two spots, with Grok 4.6 close behind at a fraction of the price. Some mirrors put GPT-5.6 Sol nearer 61 than the 58.9 shown here, which is exactly why the ordering is the signal and the decimal is not.
On individual tests, Qwen3.8-Max reports GPQA Diamond 92.6 and Muse Glimmer reports AIME 2026 94.7, both provider-reported.
Best LLM for X: decision framework
- Choose Claude Opus 5 if you want the top independent score and price is secondary to capability.
- Choose Gemini 3.7 Flash if you are coding at volume and want the cheapest frontier-adjacent rate through year-end.
- Choose Grok 4.6 if you want a near-top score at the lowest output price and can confirm the context window you need.
- Choose Meta Muse Glimmer if you need an Apache-2.0 model that runs on a single GPU with no license carve-outs.
- Choose DeepSeek V4-Pro if you need 1M context under MIT and can schedule work around off-peak pricing.
- Avoid assuming a license from “open weights.” Qwen3.8-Max ships a custom license and GLM-5.3’s flagship license is unstated, so neither is a safe Apache or MIT assumption.
Common mistakes when picking an LLM
- Trusting a single “SWE-bench #1” claim. The public sources conflict this month, and swebench.com returned no structured data, so any lone ranking is unverifiable.
- Ignoring price at scale. Sol at $20 output, Grok 4.6 at $6, and Gemini 3.7 Flash at $3.75 are a 5x spread on the same job, and that gap compounds across millions of tokens.
- Conflating open weights with a permissive license. Qwen3.8-Max is a custom license and GLM-5.3’s flagship terms are unstated, so downloadable does not mean freely shippable.
- Budgeting around a living-leaderboard decimal. The AA Index moves, and GPT-5.6 Sol already reads differently across mirrors, so re-pull the number before you commit.
- Designing around an aggregator-only context window. Grok 4.6’s 500K and Gemini 3.7 Flash’s 1M are both unconfirmed by the vendor, so verify the window before you build long-input flows.
Recent platform updates in August 2026

| Date | Event | Why it matters |
|---|---|---|
| Aug 5 | Anthropic inference hooks (beta, Enterprise) for real-time data-loss prevention | Governance for agent deployments |
| Aug 6 | Anthropic Claude Code self-hosted environments (public beta) | Enterprise on-prem coding |
| Aug 10 | Meta Muse Glimmer open weights (Apache-2.0, 30B) | Meta back in open weights with a real license |
| Aug 12 | xAI Grok 4.6 GA ($2/$6 under 200K) | New frontier model, cheapest at the top |
| Aug 12 | Alibaba Qwen3.8-Max open weights on Hugging Face | First Max-class downloadable Qwen |
| Aug 13 | Google Gemini 3.7 Flash GA | The workhorse that shipped instead of 3.5 Pro |
| Aug 12-13 | DeepSeek V4-Pro formal GA (1.6T) | Agent-focused flagship reaches GA |
| Aug 14 | Z.ai GLM-5.3 (API; weights Aug 28) | Open coding model, flagship license unstated |
| Aug 16 | DeepSeek V4 peak / off-peak pricing (16:00 UTC) | New API economics, off-peak at half |
| Aug 17 | Google Imagen 4 hard shutdown (three endpoints) | Forced migration to the Gemini image model |
| Aug 20 | Anthropic computer use, browser use, Skills API, and Files API reach GA | Agent tooling goes GA |
| Aug 21 | OpenAI GPT-5.6 Sol price cut ($5/$30 to $4/$20) | Frontier price war escalates |
| Aug 21 | Anthropic Claude Mythos 5 into Claude Security plus a $35M defender fund | Frontier security model to enterprise defenders |
| Aug 26 | Google Gemini 3.5 Transcribe GA | Speech-to-text goes GA |
| Aug 27 | Google Gemini Omni Flash GA (video, up to 4K) | Video generation goes GA |
| Aug 28 | Tencent Hunyuan Hy4 open-weight preview (770B); GLM-5.3 weights | Chinese open-weight preview at the frontier |
How to actually pick one for production
A benchmark ranks models on generic tasks, not on yours. The gap between a provider’s score and your production number is where the wrong pick hides, and it is why a model that passes evals can still fail in production. Close it with three moves.
First, build a small evaluation set from your real traffic, a few hundred prompts that look like what your users actually send.
Second, run each candidate through Future AGI with custom evals: write a grading rule, choose an LLM judge or a deterministic check, point it at your dataset columns, and set a pass/fail threshold. Each result comes back as a score with a reason, so you learn why a model failed, not just that it did.
Third, gate the winner as a CI check so a regression cannot reach production unnoticed.
For agent workloads, simulate multi-turn runs before you commit, and read the evaluation docs for metric details. The core is Apache-2.0 and self-hostable via Docker Compose, so the whole loop can run in your own environment.
Sources
- Anthropic: claude.com/blog, releasebot.io
- OpenAI: cellcog.ai, enterprisedna.co, Wikipedia GPT-5.6
- Google: blog.google, ai.google.dev changelog, SemiAnalysis
- xAI: x.ai/news, docs.x.ai release notes
- Meta: VentureBeat, MarkTechPost
- Alibaba: Hugging Face Qwen, SQ Magazine
- Z.ai: DataNorth, cellcog.ai
- Tencent: TechNode, Hugging Face Tencent
- DeepSeek: api-docs.deepseek.com, Engadget
- Benchmarks: Artificial Analysis (mirrored via benchlm.ai), SWE-bench
Previous: Best LLMs of July 2026
Frequently Asked Questions
What is the best LLM in August 2026?
What is the best LLM for coding in August 2026?
What is the best open-source LLM in August 2026?
Did Google release Gemini 3.5 Pro in August 2026?
What is the cheapest frontier LLM in August 2026?
What new LLMs launched in August 2026?
What is Meta Muse Glimmer?
Is Qwen3.8-Max open source in August 2026?
Best LLMs of June 2026 by use case: Claude Fable 5 for raw coding, GLM-5.2 for open-weight value, GPT-5.5 for agents, Gemini 3.1 Pro for long-context multimodal.
Best Voice AI June 2026: Deepgram Nova-3 for streaming STT, Cartesia Sonic-3.5 for TTS, Retell for voice agents, plus latency budgets and cost at scale.
Best LLMs May 2026: compare GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and DeepSeek V4 across coding, agents, multimodal, cost, and open weights.