Best LLMs of July 2026: Top Closed-Source, Open-Weight, Multimodal, and Coding Picks
Best LLMs of July 2026 by use case: Claude Opus 5 for agentic coding, GPT-5.6 Sol for reasoning, Kimi K3 for open-weight scale, Gemini 3.6 Flash for speed.
Table of Contents
Series note. This is the July 2026 entry in our monthly best-LLMs series. The closed frontier reshuffled in July while the open-weight tier hit a new size record. We track what shipped, what won each category, and what the public leaderboards miss. Previous: June 2026 ←. The open-weight coding price war went mainstream. May 2026 ←←. The model layer rested while infrastructure moved.

TL;DR: Best LLM per category, July 2026
| Use case | Best pick | Why | Output $/M tokens |
|---|---|---|---|
| Balanced agentic frontier (GA) | Claude Opus 5 (1M context) | New $5/$25 mid-flagship at half of Fable 5, built for agentic coding | $25 |
| Highest capability (GA) | Claude Fable 5 | 95.0% SWE-bench Verified (provider-reported), the top Claude tier | $50 |
| Reasoning (independent) | GPT-5.6 Sol | 94.1% GPQA Diamond (Artificial Analysis, independent) | $30 |
| Cheap balanced default | Claude Sonnet 5 | Balanced coding and tool-use tier, $2/$10 introductory | $10 (intro) → $15 |
| Open-weight scale leader | Kimi K3 (Modified MIT) | Largest open-weight model ever, 2.8T parameters, 1M context | $15 |
| Cost-performance open coder | DeepSeek V4-Pro (MIT) | 1.6T mixture-of-experts, 1M context, 80.6% SWE-bench Verified (provider) | ~$1.10 (aggregator) |
| Cheapest frontier-class open | DeepSeek V4-Flash-0731 (MIT) | 284B mixture-of-experts public beta for routine bulk work | $0.28 |
| Cheap fast frontier | Gemini 3.6 Flash | 17% fewer output tokens than 3.5 Flash, 1M context | $7.50 |
| High-volume subagent tier | Gemini 3.5 Flash-Lite | Cheap agentic fan-out, 1M context | $2.50 |
| Reasoning + real-time data | Grok 4.5 | 500K context, X-native grounding | $6 |
| Cheapest Apache open coder | Tencent Hunyuan Hy3 (Apache-2.0) | 295B mixture-of-experts, 256K context | $0.53 |
| Cheapest frontier tier overall | GPT-5.6 Luna | Price cut 80% on July 30 | $1.20 |
| Balanced Gemini frontier | Gemini 3.5 Flash | 92.2% GPQA Diamond (Artificial Analysis) at $9 output | $9.00 |
If you only read one row: Claude Opus 5 for balanced agentic coding at half Fable 5’s price, Kimi K3 for open-weight scale, GPT-5.6 Sol for independently verified reasoning, DeepSeek V4-Flash-0731 for the cheapest frontier-class open weights, and Gemini 3.6 Flash for cheap high-throughput work. Everything else is a trade-off around those five.

The single biggest story of July 2026: the closed frontier reshuffled while open weights hit a new scale ceiling. Anthropic shipped Claude Opus 5 (July 24) as a new $5/$25 middle flagship at roughly half Fable 5’s price, and OpenAI completed the GPT-5.6 rollout to general availability (July 9) then cut Terra 20% and Luna 80% on July 30.
In the same window, Moonshot’s Kimi K3 (open weights July 27, 2.8T parameters, Modified MIT) became the largest open-weight model ever shipped. The frontier is now a pricing decision at the top and a scale story at the bottom, moving in the same month.
The story of July 2026: a cheaper mid-flagship, tier-wide price cuts, and a new open-weight record
June was the open-weight coding price war. July was the closed frontier repricing itself while Chinese labs set a size record. Three things happened at once, and each one changes a production decision.
Anthropic opened the month’s headline. Claude Opus 5 went generally available July 24 at $5 input / $25 output per million tokens, with a 1M-token context window and 128k max output.
That is half the price of Claude Fable 5, and Anthropic positions it for complex agentic coding and enterprise work rather than a single headline benchmark. Fable 5 stays the capability ceiling at $10/$50, so the Claude lineup now reads Opus 5 for balanced production, Fable 5 for the hardest jobs, and Sonnet 5 for cheap volume.
OpenAI moved on price. The GPT-5.6 family (Sol, Terra, and Luna) reached general availability July 9, and on July 30 OpenAI cut Terra 20% to $2/$12 and Luna 80% to $0.20/$1.20, while Sol held at $5/$30. GPT-5.6 Sol leads independently measured reasoning at 94.1% GPQA Diamond on Artificial Analysis, so the flagship kept its price and the cheaper tiers absorbed the cut.

Then the scale record. Moonshot shipped Kimi K3 to open weights July 27 under a Modified MIT license, 2.8 trillion total parameters and a 1M-token context window, the largest open-weight model ever released. It is not cheap for an open model at $3/$15, because a 1.56TB download and a huge active-parameter count carry real inference cost.
DeepSeek went the other way with V4-Pro (generally available July 20, MIT, 1.6T mixture-of-experts) and the V4-Flash-0731 public beta July 31 at $0.28 output, and Tencent shipped Hunyuan Hy3 July 6 under Apache-2.0 at about $0.53 output.
Google shipped three Gemini models on July 21, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite among them, but not Gemini 3.5 Pro, which remained in partner testing. xAI released Grok 4.5 to the public July 8 with a 500K context window and real-time data from X. Alibaba previewed Qwen3.8-Max on July 19 as an API-only model, so its flagship Max tier stays closed while smaller Qwen sizes remain open.
The takeaway from July 2026: the top of the market got cheaper and the open tier got larger, and both moves reward the same discipline. Model choice is a price, license, and reliability decision, and the gap between a provider benchmark and your production number is what decides whether the cheap pick or the capable one actually holds.
Top closed-source / proprietary LLMs in July 2026
Claude Opus 5. Best for balanced agentic coding
Anthropic. Generally available July 24, 2026. 1M context, 128k max output.
The July headline, and the new default for most Anthropic production work. Opus 5 lands at $5 input / $25 output per million tokens, half the price of Fable 5, and Anthropic describes it as built for complex agentic coding and enterprise work. It keeps the 1M-token context window and 128k max output of the top Claude tier.
On the independent Artificial Analysis Intelligence Index it is narrowly the top composite model at 61, effectively tied with Fable 5 at 60 and ahead of GPT-5.6 Sol at 59.
- 1M-token context window, 128k max output
- $5 input / $25 output per million tokens
- Positioned for agentic coding and enterprise work
- Best for: production coding agents and long-horizon enterprise work that wants frontier capability without the Fable 5 price
What it does not win: the capability ceiling (Fable 5 sits above it), and rock-bottom cost (open-weight coders undercut it on output tokens). Anthropic frames Opus 5 on agentic-coding benchmarks rather than a single SWE-bench headline, so reproduce it on your own tasks before you commit to a number.
Claude Fable 5. Best for highest capability
Anthropic. Generally available June 9, 2026. 1M context, 128k max output. Carryover into July.
The capability ceiling of the Claude lineup, now the priciest tier at $10 input / $50 output per million tokens. Fable 5 posts 95.0% on SWE-bench Verified as a provider-reported figure, and its GPQA Diamond of 92.6% on Artificial Analysis trails Opus 4.8, GPT-5.6 Sol, and Gemini 3.1 Pro, so read it as a coding-first model rather than a reasoning leader.
- 1M-token context window, 128k max output
- $10 input / $50 output per million tokens
- 95.0% SWE-bench Verified (provider-reported)
- Best for: the hardest multi-file coding and reasoning work where capability outranks cost
What it does not win: cost (it is the most expensive tier here), and general reasoning headlines (its GPQA Diamond trails the reasoning leaders).
Claude Sonnet 5. Best cheap balanced tier
Anthropic. Generally available June 30, 2026. Carryover into July.
The cheap, balanced coding and tool-use tier below Opus 5. Sonnet 5 launched at $2 input / $10 output introductory pricing through August 31, 2026, then $3/$15, with the same 1M-token context window and 128k max output as the larger Claude models.
- 1M-token context window, 128k max output
- $2 input / $10 output introductory through August 31, 2026, then $3 / $15
- Balanced coding and tool-use tier
- Best for: high-volume agents, chat surfaces, and cost-sensitive production where Opus 5 is overkill
What it does not win: top-end capability (Opus 5 and Fable 5 sit above it), and rock-bottom price (open-weight coders undercut it on output tokens).
GPT-5.6 Sol. Best for independently verified reasoning
OpenAI. Generally available July 9, 2026. Native vision and audio.
OpenAI’s flagship for reasoning-heavy work, and the July pick when you want an independently measured score. Sol posts 94.1% on GPQA Diamond per Artificial Analysis, an independent tracker, and holds the top Artificial Analysis Intelligence Index position among the GPT-5.6 tiers. Its price held at $5/$30 through the July 30 cuts that hit the cheaper tiers.
- $5 input / $30 output per million tokens (unchanged July 30)
- 94.1% GPQA Diamond (Artificial Analysis, independent)
- Native vision and audio
- Best for: reasoning-heavy applications, agentic work that wants a hosted flagship, broad ecosystem support
What it does not win: cost (Terra and Luna undercut it heavily after the July 30 cuts), and open-weight licensing (it is hosted only). OpenAI published no SWE-bench Verified figure at launch, so pair it with your own coding eval.
GPT-5.6 Terra and Luna. The price-cut mid and low tiers
OpenAI. Generally available July 9, 2026. Repriced July 30, 2026.
The two cheaper GPT-5.6 tiers, and where OpenAI spent its July pricing move. On July 30, Terra dropped 20% to $2/$12 and Luna dropped 80% to $0.20/$1.20, which makes Luna one of the cheapest hosted frontier-family tiers available. Both carry Artificial Analysis Intelligence Index positions below Sol (Terra 55, Luna 51 on the 100-point scale).
- Terra: $2 input / $12 output per million tokens (cut 20% July 30 from $2.50/$15)
- Luna: $0.20 input / $1.20 output per million tokens (cut 80% July 30 from $1/$6)
- Both native vision and audio
- Best for: high-volume production (Terra) and cost-sensitive bulk work (Luna) inside the OpenAI ecosystem
What it does not win: top-end reasoning (Sol leads the family), and the open-weight economics of the DeepSeek and Tencent tiers on the cheapest bulk work.
Gemini 3.6 Flash. Best cheap fast frontier
Google DeepMind. Generally available July 21, 2026.
The cheap, fast frontier option, and Google’s July speed release. Gemini 3.6 Flash runs at $1.50 input / $7.50 output per million tokens with a 1,048,576-token context window, and Google reports it using 17% fewer output tokens than 3.5 Flash for the same work, which lowers real cost below the sticker gap. Provider benchmarks put it at 58.7% on SWE-bench Pro and 83% on OSWorld Verified.
- $1.50 input / $7.50 output per million tokens
- 1,048,576-token context window, 65,536 max output
- 17% fewer output tokens than 3.5 Flash (provider-reported)
- Best for: high-volume, latency-sensitive agentic work that wants frontier-adjacent quality cheaply
What it does not win: top-end coding and reasoning (the Pro tiers and Opus 5 lead), and open-weight economics for the cheapest bulk output.
Grok 4.5. Best for reasoning with real-time data
xAI. Public July 8, 2026. 500K context. Proprietary, API-only.
xAI’s current flagship, and the pick when reasoning meets a need for current-events grounding. Grok 4.5 ships proprietary and API-only with a 500K-token context window (down from Grok 4.3’s 1M) at $2 input / $6 output, with prompt caching available. It pulls real-time data from X without a separate retrieval layer.
- $2 input / $6 output per million tokens
- 500K-token context window
- Real-time data via X
- Best for: research agents that need current data, reasoning-heavy workflows, X-native grounding
What it does not win: open-weight licensing (it is hosted only), and the longest context window (it dropped to 500K). Its coding and terminal benchmarks are single-source at launch, so verify them before you rely on them.
Top open-weight and Chinese-frontier LLMs in July 2026
The open-weight tier is where July set a record. Kimi K3 became the largest open-weight model ever released, DeepSeek and Tencent held the cost-performance floor, and Alibaba kept its flagship Max tier closed. Weigh license and real inference cost alongside the download size, because the biggest model is not the cheapest to run.
Kimi K3. The July open-weight record
Moonshot AI. API July 16, open weights July 27, 2026. Modified MIT (“Kimi K3 License”), 1M context.
The single most important open-weight release of July 2026, and the largest open-weight model ever shipped. Kimi K3 carries 2.8 trillion total parameters and a 1M-token context window under a Modified MIT license (Moonshot’s own “Kimi K3 License”), with a download near 1.56TB.
The license is permissive for most users but gates commercial use at scale: teams above $20M in annual model-as-a-service revenue or 100M monthly active users need a separate agreement with Moonshot and must display “Kimi K3” attribution.
Moonshot reports strong coding and reasoning numbers (GPQA Diamond 93.5%, SWE-bench Verified 76.8%, Terminal-Bench 2.1 88.3%), all provider-reported, so treat them as claims to reproduce.
The price reflects the size. Kimi K3 lists $3 input / $15 output per million tokens, which is expensive for an open model because a 2.8T-parameter network with a large active-parameter count is costly to serve. The value here is open weights at record scale, not the cheapest output.
- Modified MIT license (the “Kimi K3 License”; large-scale commercial use over $20M revenue or 100M MAU needs a separate agreement), downloadable weights (~1.56TB)
- 2.8T total parameters, 1M-token context window
- Provider-reported GPQA Diamond 93.5%, SWE-bench Verified 76.8% (independent runs pending)
- $3 input / $15 output per million tokens
Best for: teams that need open weights at frontier scale and can run or rent the inference, and who will benchmark the provider numbers on their own domain.
Skip if: you need an independently verified score today, or output price is your dominant constraint (DeepSeek and Tencent undercut it heavily).
DeepSeek V4-Pro and V4-Flash-0731. Best cost-performance on the open frontier
DeepSeek. V4-Pro generally available July 20; V4-Flash-0731 public beta July 31, 2026. MIT license.
The cost-performance anchors of the open tier, both MIT and both current in July. V4-Pro is the coder at 80.6% SWE-bench Verified (provider) and 91.6% LiveCodeBench (provider), a 1.6T mixture-of-experts with a 1M-token context window. V4-Flash-0731 is the routine-work beta at a fraction of the price.
- V4-Pro: MIT, 1.6T / 49B active, 1M context; 80.6% SWE-bench Verified (provider); API around $0.27 / $1.10 (aggregator-reported, verify on the provider page)
- V4-Flash-0731: MIT, 284B / 13B active, 1M context; $0.14 input / $0.28 output per million tokens
- Both downloadable weights
Best for: high-volume coding and general work where output price dominates, and teams that can self-host or accept inference from DeepSeek’s API.
Skip if: you need an independent benchmark today (the published scores are provider-reported), or you want a single hosted flagship with vendor support.
Tencent Hunyuan Hy3. Cheapest permissive coder
Tencent. Generally available July 6, 2026. Apache-2.0, 256K context.
Tencent’s July open release, and the cheapest permissively licensed coder on this list. Hunyuan Hy3 is a 295B mixture-of-experts (192 experts, top-8 routing, 21B active) with a 256K-token context window, listed around $0.13 input / $0.53 output on OpenRouter. Its provider SWE-bench Verified of 74.4% is single-source, so name it and verify before you rely on it.
- Apache-2.0 license, 256K-token context window
- 295B total / 21B active mixture-of-experts
- ~$0.13 input / $0.53 output per million tokens (aggregator-reported)
- 74.4% SWE-bench Verified (provider, verify)
Best for: cost-sensitive coding on a truly permissive license, and teams that want an Apache-2.0 alternative to the MIT Chinese models.
Skip if: you need a published independent score, or a context window longer than 256K for whole-repository work.
GLM-5.2. The June headliner, still current
Z.ai / Zhipu. Generally available June 13, 2026. MIT license, 1M context. No July update.
June’s price-war headliner, carried into July unchanged. GLM-5.2 ships under an MIT license with a 1M-token context window and roughly 750B total parameters, and Z.ai did not ship a GLM-5.3 or GLM-5.5 in July, so this is the current Z.ai open flagship. Its coding claims are vendor-reported, as covered in the June edition.
- MIT license, downloadable weights
- ~750B total / ~40B active, 1M-token context window
- No July version update
Best for: self-hosted or hosted coding agents that want an MIT license and a 1M context, with your own eval to confirm the vendor coding numbers.
Skip if: you need the newest open release (Kimi K3, DeepSeek V4-Pro, and Tencent Hy3 all shipped in July).
Qwen3.8-Max, Mistral Large 3, and Llama 4. The closed preview and the carryovers
Alibaba (preview July 19, 2026), Mistral AI (December 2, 2025), and Meta (April 2025).
Three names worth tracking, with different availability. Qwen3.8-Max previewed July 19 as an API-only model, so Alibaba’s flagship Max tier stays closed at launch while smaller Qwen sizes remain open; do not treat it as open-weight. Mistral Large 3 remains the strongest non-Chinese permissive option at Apache-2.0, 675B mixture-of-experts, and a 262K context window. Llama 4 Scout and Maverick carry unchanged from April 2025 under the Llama Community License.
- Qwen3.8-Max: API-only preview, closed weights (not an open-weight pick)
- Mistral Large 3: Apache-2.0, 675B / 41B active, 262K context
- Llama 4 Scout / Maverick: 10M / 1M context, Llama Community License
Best for: a truly permissive non-Chinese license (Mistral Large 3), and long-context open deployments (Llama 4 Scout at 10M tokens).
For formal-math work specifically, Mistral also shipped Leanstral 1.5 on July 2 (Apache-2.0, 119B / 6B active, 256K context), a Lean 4 proof agent that posts 587/672 on PutnamBench. Treat it as a domain specialist, not a general LLM pick.
Top multimodal LLMs in July 2026
Multimodal is three separate races now: vision, image generation, and video generation, each with its own leaders and price floors. Two deprecations matter this month, so read the notes below before you build.
Vision (image and document understanding)
| Use case | Best closed | Best open |
|---|---|---|
| General image understanding | Gemini 3.1 Pro | Qwen open VL family |
| Document / OCR / chart reading | Claude Opus 5 / Fable 5 | Qwen open VL family |
| Long-context multi-image | Gemini 3.6 Flash (1M tokens) | Llama 4 Scout (10M) |
| Screen / UI understanding | GPT-5.6 Sol | Qwen open VL family |
Every Western frontier model now handles vision natively, including Opus 5, Fable 5, GPT-5.6, Gemini 3.6 Flash, and Grok 4.5. The open column has closed the gap on the basics while still trailing on dense charts and technical diagrams, so confirm any open VL model on your own document distribution before you switch.
Image generation
The image-gen category stayed fragmented in July, and one flagship left the market. Pick by what the image is for.
| Use case | Best pick | Why |
|---|---|---|
| General-purpose closed | GPT Image 2 | Still the OpenAI image flagship |
| Google flagship image | Gemini 3 Pro Image | Top Gemini image tier |
| Fast Google image | Gemini 3.1 Flash Image | Speed-tuned Gemini image model |
| Cheap image tier (new) | Gemini 3.1 Flash-Lite Image | New July 1, $0.034 per 1,000 images, ~4s |
| Open-weight / self-host | FLUX.2 Pro | On your hardware, from Black Forest Labs |
| Alternative closed | Seedream v5 | ByteDance image model |
Two July notes. Google added Gemini 3.1 Flash-Lite Image (Nano Banana 2 Lite) on July 1 at $0.034 per 1,000 images. Black Forest Labs announced FLUX 3 on July 23, but only the video variant is in early access, so treat FLUX 3 image as announced, not shipped. Drop Imagen 4 and Imagen 4 Ultra entirely: they were deprecated June 15 with a hard shutdown on August 17, 2026.
Video generation
Video-gen leadership held through July, with one clear deprecation to plan around.
| Use case | Best pick | Why |
|---|---|---|
| Best all-around | Google Veo 3.1 | Leads prompt adherence, native audio, 4K output |
| Multi-shot storytelling | Kling 3.0 | Sequences with subject consistency across cuts |
| Native audio-video (preview) | Gemini Omni Flash | Developer API preview since June 30 |
For production video in July 2026, Veo 3.1 is the default for narrative scenes and Kling 3.0 for multi-shot work. ByteDance shipped Seedance 2.5 on July 31, but only on its consumer app, so the API is not live yet; hold it until the developer endpoint ships. Remove OpenAI Sora 2 from new builds: the web and app went dark April 26, and the API ends September 24, 2026.
Audio understanding
Native audio handling (speech as input, without a separate transcription step) is standard across the Western frontier models. Google’s Gemini Omni Flash adds native audio in developer preview, and GPT-5.6 and Gemini 3.1 Pro both handle speech-in directly, which preserves prosody and disambiguates accents that break transcribe-then-prompt pipelines. For the full speech stack, see the dedicated voice guide below.
Voice and audio (covered separately)
Voice AI has its own decision logic. STT (AssemblyAI, Deepgram, ElevenLabs, OpenAI), TTS (Cartesia, ElevenLabs, Deepgram, Hume), and voice-agent platforms (LiveKit, Pipecat, Vapi, Retell) follow different picks than text-only LLMs. The conversational latency budget alone, anchored on the ITU-T G.114 one-way delay recommendation, drives different choices than text-LLM workflows.
See Best Voice AI of July 2026 for the full STT, TTS, and voice-agent stack.
Top embeddings and retrieval models in July 2026
Embeddings have two production constraints: retrieval quality on your domain, and price per million tokens. The public leaderboards churn month to month, so the honest move is to name the durable options and run your own retrieval eval rather than chase a live board position.
| Use case | Durable pick | Notes |
|---|---|---|
| Contextualized retrieval (new) | voyage-context-4 | Chunk embeddings with document context, $0.12 / 1M |
| Best general retrieval (closed) | voyage-4-large | Mixture-of-experts, Voyage’s top retrieval model |
| Multimodal retrieval | gemini-embedding-001 | Text plus image, audio, and PDF in one model |
| OpenAI ecosystem default | text-embedding-3-large | Strong general default; -3-small for cheapest viable |
| Multilingual enterprise | Cohere embed-v4 | Multilingual production coverage |
| Open-weight / self-host | Qwen3-Embedding-8B | Strong open multilingual option on your hardware |
Voyage’s current pick is voyage-context-4 (June 29, contextualized chunk embeddings at $0.12 per million tokens) or voyage-4-large; voyage-3-large is superseded, so do not cite it as current.
We are deliberately not printing MTEB numbers here, because the board positions move often and several current figures did not clear a two-source check at publication. Shortlist two or three of the models above, score them on your own documents, and add a reranker for a few points of retrieval quality at small added cost.
Top coding-specific LLMs in July 2026
Coding is the highest-stakes category, and July pushed both ends: a cheaper capable pick at the top (Opus 5) and a new scale leader in open weights (Kimi K3). One caveat frames the whole table: independent SWE-bench Verified rankings were not confirmable in July, so every score below is provider-reported unless labeled otherwise.
| Use case | Top pick | Score | Output $/M tokens |
|---|---|---|---|
| Balanced agentic coding | Claude Opus 5 | Agentic-coding focus (no single provider SWE headline) | $25 |
| Raw capability | Claude Fable 5 | 95.0% SWE-bench Verified (provider) | $50 |
| Open-weight scale | Kimi K3 | 76.8% SWE-bench Verified (provider) | $15 |
| Cost-performance open coder | DeepSeek V4-Pro | 80.6% SWE-bench Verified (provider) | ~$1.10 |
| Cheapest Apache open coder | Tencent Hunyuan Hy3 | 74.4% SWE-bench Verified (provider) | $0.53 |
| Cheap fast frontier coding | Gemini 3.6 Flash | 58.7% SWE-bench Pro (provider) | $7.50 |
| Cheapest frontier-class open | DeepSeek V4-Flash-0731 | Provider harness only | $0.28 |
One label matters here. The independent SWE-bench Verified leaderboard did not return structured data at publication, and only Terminal-Bench 2.0 was fetchable, so we did not print an independent coding ranking for July. Every score above is provider-reported. The gap between a provider score and your reproduction is the number that decides your pick, so run a domain eval on your own repositories before you trust any single figure.
Top reasoning and math LLMs in July 2026
Reasoning is where July’s cleanest independent numbers live, because GPQA Diamond has an Artificial Analysis run that clears a two-source check. GPT-5.6 Sol leads at 94.1% GPQA Diamond (Artificial Analysis, independent), with Gemini 3.5 Flash at 92.2% and Claude Fable 5 at 92.6% on the same board.
On provider-reported reasoning, Gemini 3.1 Pro’s model card lists GPQA Diamond 94.3%, SWE-bench Verified 80.6%, and HLE 44.4% without tools, and Claude Opus 4.8’s system card lists GPQA Diamond 93.6%. Claude Opus 5’s GPQA reports span 84.1% to 93.7% depending on the effort setting, so we name the model and skip a single figure.
Meta’s closed Muse Spark posts HLE 58% in its highest-effort mode (provider). If reasoning is your core workload, weight the independent GPQA Diamond result above any single-source claim and run your own eval on a graduate-level set that matches your domain.
Best LLM for X: decision framework
Choose Claude Opus 5 if:
- You want balanced frontier capability for agentic coding at half the Fable 5 price.
- You need a 1M-token context window for long-horizon or whole-repository work.
- You are on the Anthropic ecosystem and want the new production default.
Choose Claude Fable 5 if:
- You need the highest available coding capability and cost is secondary.
- Your work is hard multi-file reasoning where the capability ceiling pays for itself.
Choose GPT-5.6 Sol if:
- Reasoning is your core workload and you want an independently measured score.
- You want a hosted flagship with native vision, audio, and broad ecosystem support.
Choose Kimi K3 if:
- You need open weights at frontier scale and can run or rent the inference.
- A Modified MIT license and a 1M context matter more than the cheapest output price.
Choose DeepSeek V4-Pro or V4-Flash-0731 if:
- Output price is the dominant constraint.
- Your workload is primarily coding (V4-Pro) or high-volume routine work (V4-Flash-0731).
- MIT licensing matters for downstream redistribution.
Choose Gemini 3.6 Flash if:
- You want cheap, fast frontier-adjacent quality at $7.50 output.
- A 1M-token context window and lower output-token counts cut your real cost.
Choose Grok 4.5 if:
- Graduate-level reasoning and real-time data are both core.
- You want X-native grounding without a separate retrieval layer.
Avoid Gemini 3.5 Pro, Qwen3.8-Max, FLUX 3 Image, and Seedance 2.5 for now. Gemini 3.5 Pro did not ship in July, Qwen3.8-Max is an API-only preview, FLUX 3’s image variant is announced but not released, and Seedance 2.5’s API is not live. None is a July production option.
Common mistakes when picking an LLM in July 2026
The four most expensive errors we see production teams make:
- Treating provider scores as independent. July’s SWE-bench Verified numbers for Kimi K3, DeepSeek V4-Pro, and Tencent Hy3 are provider-reported, and the independent leaderboard did not render. Confirm coding scores on your own repositories before you migrate.
- Assuming the biggest open model is the cheapest. Kimi K3 is the largest open-weight model ever, and at $15 output it costs more than several Western mid-tiers. Scale is not savings, so price the inference before you adopt it.
- Ignoring total cost of ownership. Listed price times token volume is the sticker number. Real cost includes retry rate, thinking and effort multipliers, and failure recovery, and a flaky agent triples the bill before you notice.
- Building on a preview as if it were generally available. Qwen3.8-Max, FLUX 3 Image, Seedance 2.5, and Gemini 3.5 Pro are not shipped production options in July. Build on generally available models and keep a fallback in any path that cannot afford downtime.
Recent platform updates
| Date | Event | Why it matters |
|---|---|---|
| July 1 | Claude Mythos 5 access restored for US orgs; Gemini 3.1 Flash-Lite Image | Restricted Fable-class access returns; new cheap image tier |
| July 2 | Mistral Leanstral 1.5 (Apache-2.0) | Lean 4 formal-math specialist |
| July 6 | Tencent Hunyuan Hy3 (Apache-2.0) | Cheapest permissive open coder |
| July 8 | xAI Grok 4.5 public | Reasoning with real-time X data, 500K context |
| July 9 | OpenAI GPT-5.6 (Sol/Terra/Luna) generally available | Flagship family reaches GA |
| July 16 | Moonshot Kimi K3 API | The record open-weight model goes live |
| July 19 | Alibaba Qwen3.8-Max preview (API-only) | Flagship Max tier stays closed |
| July 20 | DeepSeek V4-Pro generally available (MIT) | Cost-performance open coder |
| July 21 | Google Gemini 3.6 Flash + 3.5 Flash-Lite (no 3.5 Pro) | Cheap fast frontier; 3.5 Pro slips |
| July 23 | FLUX 3 announced (early access, video only) | Next FLUX generation, image not yet shipped |
| July 24 | Anthropic Claude Opus 5 generally available | New $5/$25 mid-flagship at half Fable 5 |
| July 27 | Kimi K3 open weights | Largest open-weight model ever released |
| July 30 | GPT-5.6 Terra/Luna price cuts | Terra down 20%, Luna down 80% |
| July 31 | DeepSeek V4-Flash-0731 beta; Seedance 2.5 (consumer) | Cheapest open beta; video model on consumer app only |
How to actually pick one for production
The leaderboard is the wrong artifact to make a July 2026 production decision from. Three things to do instead:
- Run a domain reproduction. Take 100 to 500 of your actual production prompts, run them through your two or three candidate models with your harness, and score them with Future AGI evals or your own judge. For July specifically, this is how you separate Opus 5’s agentic-coding framing and Kimi K3’s provider scores from the number you will actually ship.
- Measure reliability under load. A public score is an aggregate over a fixed task set, and it does not predict variance across your prompts and repeated agent runs. Use Future AGI Simulate to stress-test agents across long sessions, where most frontier models lose a chunk of headline accuracy that benchmarks never surface.
- Cost-adjust every score. July’s spread runs from $0.28 to $50 per million output tokens. Compute score-per-dollar on your domain before you default to a brand, because Kimi K3’s record scale and DeepSeek’s cheap output only pay off if they hold up on your traffic.
July rewarded teams that treated model choice as a price, license, and reliability decision. The top of the market got cheaper with Opus 5 and the GPT-5.6 cuts, the open tier got larger with Kimi K3, and both moves reward the team that measures before it commits.
The practical answer is to shortlist by your dominant workload, then let your own reproduction and reliability numbers pick the winner. Put more effort into the eval loop above the model than into comparing leaderboard rows, because that is the layer that decides whether July’s cheapest pick or its most capable one actually works for you.
Sources
Frontier model launches (primary):
- Claude models overview (Anthropic)
- Claude Opus 5 (Anthropic)
- Anthropic pricing
- OpenAI GPT-5.6
- Gemini API changelog (Google)
- Gemini API pricing (Google)
- DeepSeek API updates
- Grok 4.5 (TechCrunch)
Open-weight and benchmarks:
- Kimi K3 (MLQ)
- Tencent Hunyuan Hy3
- GLM-5.2 analysis (Interconnects)
- Artificial Analysis Intelligence Index
- SWE-bench Verified leaderboard
- GPQA Diamond via Artificial Analysis
Previous: Best LLMs of June 2026 ←. The open-weight coding price war went mainstream and Anthropic opened a new top tier.
Frequently Asked Questions
What is the best LLM in July 2026?
What is the best open-source LLM in July 2026?
What is the best LLM for coding in July 2026?
What new LLMs were released in July 2026?
What is Claude Opus 5?
What is Kimi K3?
How much does Claude Opus 5 cost?
Did OpenAI cut GPT-5.6 prices in July 2026?
Best LLMs of June 2026 by use case: Claude Fable 5 for raw coding, GLM-5.2 for open-weight value, GPT-5.5 for agents, Gemini 3.1 Pro for long-context multimodal.
Best LLMs May 2026: compare GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and DeepSeek V4 across coding, agents, multimodal, cost, and open weights.
Best LLMs April 2026: compare GPT-5.5, Claude Opus 4.7, DeepSeek V4, Gemma 4, and Qwen after benchmark trust broke and prices compressed fast.