Research

Best LLMs of July 2026: Top Closed-Source, Open-Weight, Multimodal, and Coding Picks

Best LLMs of July 2026 by use case: Claude Opus 5 for agentic coding, GPT-5.6 Sol for reasoning, Kimi K3 for open-weight scale, Gemini 3.6 Flash for speed.

· 26 min read
best-llms monthly-compare 2026 frontier-models open-source multimodal coding model-comparison
Two-cluster map of the large language models that shipped in July 2026, a closed frontier group and an open-weight frontier group of white nodes on black, with Claude Opus 5 the brightest node as the month's new mid-flagship.
Table of Contents

Series note. This is the July 2026 entry in our monthly best-LLMs series. The closed frontier reshuffled in July while the open-weight tier hit a new size record. We track what shipped, what won each category, and what the public leaderboards miss. Previous: June 2026 ←. The open-weight coding price war went mainstream. May 2026 ←←. The model layer rested while infrastructure moved.

Two-cluster map of the large language models that shipped in July 2026, a closed frontier group and an open-weight frontier group of white nodes joined by faint blueprint lines on black, with Claude Opus 5 the brightest node as the month's new mid-flagship.

TL;DR: Best LLM per category, July 2026

Use caseBest pickWhyOutput $/M tokens
Balanced agentic frontier (GA)Claude Opus 5 (1M context)New $5/$25 mid-flagship at half of Fable 5, built for agentic coding$25
Highest capability (GA)Claude Fable 595.0% SWE-bench Verified (provider-reported), the top Claude tier$50
Reasoning (independent)GPT-5.6 Sol94.1% GPQA Diamond (Artificial Analysis, independent)$30
Cheap balanced defaultClaude Sonnet 5Balanced coding and tool-use tier, $2/$10 introductory$10 (intro) → $15
Open-weight scale leaderKimi K3 (Modified MIT)Largest open-weight model ever, 2.8T parameters, 1M context$15
Cost-performance open coderDeepSeek V4-Pro (MIT)1.6T mixture-of-experts, 1M context, 80.6% SWE-bench Verified (provider)~$1.10 (aggregator)
Cheapest frontier-class openDeepSeek V4-Flash-0731 (MIT)284B mixture-of-experts public beta for routine bulk work$0.28
Cheap fast frontierGemini 3.6 Flash17% fewer output tokens than 3.5 Flash, 1M context$7.50
High-volume subagent tierGemini 3.5 Flash-LiteCheap agentic fan-out, 1M context$2.50
Reasoning + real-time dataGrok 4.5500K context, X-native grounding$6
Cheapest Apache open coderTencent Hunyuan Hy3 (Apache-2.0)295B mixture-of-experts, 256K context$0.53
Cheapest frontier tier overallGPT-5.6 LunaPrice cut 80% on July 30$1.20
Balanced Gemini frontierGemini 3.5 Flash92.2% GPQA Diamond (Artificial Analysis) at $9 output$9.00

If you only read one row: Claude Opus 5 for balanced agentic coding at half Fable 5’s price, Kimi K3 for open-weight scale, GPT-5.6 Sol for independently verified reasoning, DeepSeek V4-Flash-0731 for the cheapest frontier-class open weights, and Gemini 3.6 Flash for cheap high-throughput work. Everything else is a trade-off around those five.

Artificial Analysis Intelligence Index for July 2026: Claude Opus 5 leads the composite score at 61, narrowly ahead of Claude Fable 5 at 60 and GPT-5.6 Sol at 59, with Kimi K3 the top open-weight model at 57, shown as a horizontal bar scoreboard with the Opus 5 bar bright white and the rest in gray.

The single biggest story of July 2026: the closed frontier reshuffled while open weights hit a new scale ceiling. Anthropic shipped Claude Opus 5 (July 24) as a new $5/$25 middle flagship at roughly half Fable 5’s price, and OpenAI completed the GPT-5.6 rollout to general availability (July 9) then cut Terra 20% and Luna 80% on July 30.

In the same window, Moonshot’s Kimi K3 (open weights July 27, 2.8T parameters, Modified MIT) became the largest open-weight model ever shipped. The frontier is now a pricing decision at the top and a scale story at the bottom, moving in the same month.

The story of July 2026: a cheaper mid-flagship, tier-wide price cuts, and a new open-weight record

June was the open-weight coding price war. July was the closed frontier repricing itself while Chinese labs set a size record. Three things happened at once, and each one changes a production decision.

Anthropic opened the month’s headline. Claude Opus 5 went generally available July 24 at $5 input / $25 output per million tokens, with a 1M-token context window and 128k max output.

That is half the price of Claude Fable 5, and Anthropic positions it for complex agentic coding and enterprise work rather than a single headline benchmark. Fable 5 stays the capability ceiling at $10/$50, so the Claude lineup now reads Opus 5 for balanced production, Fable 5 for the hardest jobs, and Sonnet 5 for cheap volume.

OpenAI moved on price. The GPT-5.6 family (Sol, Terra, and Luna) reached general availability July 9, and on July 30 OpenAI cut Terra 20% to $2/$12 and Luna 80% to $0.20/$1.20, while Sol held at $5/$30. GPT-5.6 Sol leads independently measured reasoning at 94.1% GPQA Diamond on Artificial Analysis, so the flagship kept its price and the cheaper tiers absorbed the cut.

Output price per million tokens for July 2026: Western closed frontier bars for Claude Fable 5 at $50, GPT-5.6 Sol at $30, and Claude Opus 5 at $25 tower over an open and Chinese frontier group where Kimi K3 sits at $15, Tencent Hunyuan Hy3 at $0.53, and DeepSeek V4-Flash-0731 at $0.28, the cheapest bar highlighted bright white.

Then the scale record. Moonshot shipped Kimi K3 to open weights July 27 under a Modified MIT license, 2.8 trillion total parameters and a 1M-token context window, the largest open-weight model ever released. It is not cheap for an open model at $3/$15, because a 1.56TB download and a huge active-parameter count carry real inference cost.

DeepSeek went the other way with V4-Pro (generally available July 20, MIT, 1.6T mixture-of-experts) and the V4-Flash-0731 public beta July 31 at $0.28 output, and Tencent shipped Hunyuan Hy3 July 6 under Apache-2.0 at about $0.53 output.

Google shipped three Gemini models on July 21, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite among them, but not Gemini 3.5 Pro, which remained in partner testing. xAI released Grok 4.5 to the public July 8 with a 500K context window and real-time data from X. Alibaba previewed Qwen3.8-Max on July 19 as an API-only model, so its flagship Max tier stays closed while smaller Qwen sizes remain open.

The takeaway from July 2026: the top of the market got cheaper and the open tier got larger, and both moves reward the same discipline. Model choice is a price, license, and reliability decision, and the gap between a provider benchmark and your production number is what decides whether the cheap pick or the capable one actually holds.

Top closed-source / proprietary LLMs in July 2026

Claude Opus 5. Best for balanced agentic coding

Anthropic. Generally available July 24, 2026. 1M context, 128k max output.

The July headline, and the new default for most Anthropic production work. Opus 5 lands at $5 input / $25 output per million tokens, half the price of Fable 5, and Anthropic describes it as built for complex agentic coding and enterprise work. It keeps the 1M-token context window and 128k max output of the top Claude tier.

On the independent Artificial Analysis Intelligence Index it is narrowly the top composite model at 61, effectively tied with Fable 5 at 60 and ahead of GPT-5.6 Sol at 59.

  • 1M-token context window, 128k max output
  • $5 input / $25 output per million tokens
  • Positioned for agentic coding and enterprise work
  • Best for: production coding agents and long-horizon enterprise work that wants frontier capability without the Fable 5 price

What it does not win: the capability ceiling (Fable 5 sits above it), and rock-bottom cost (open-weight coders undercut it on output tokens). Anthropic frames Opus 5 on agentic-coding benchmarks rather than a single SWE-bench headline, so reproduce it on your own tasks before you commit to a number.

Claude Fable 5. Best for highest capability

Anthropic. Generally available June 9, 2026. 1M context, 128k max output. Carryover into July.

The capability ceiling of the Claude lineup, now the priciest tier at $10 input / $50 output per million tokens. Fable 5 posts 95.0% on SWE-bench Verified as a provider-reported figure, and its GPQA Diamond of 92.6% on Artificial Analysis trails Opus 4.8, GPT-5.6 Sol, and Gemini 3.1 Pro, so read it as a coding-first model rather than a reasoning leader.

  • 1M-token context window, 128k max output
  • $10 input / $50 output per million tokens
  • 95.0% SWE-bench Verified (provider-reported)
  • Best for: the hardest multi-file coding and reasoning work where capability outranks cost

What it does not win: cost (it is the most expensive tier here), and general reasoning headlines (its GPQA Diamond trails the reasoning leaders).

Claude Sonnet 5. Best cheap balanced tier

Anthropic. Generally available June 30, 2026. Carryover into July.

The cheap, balanced coding and tool-use tier below Opus 5. Sonnet 5 launched at $2 input / $10 output introductory pricing through August 31, 2026, then $3/$15, with the same 1M-token context window and 128k max output as the larger Claude models.

  • 1M-token context window, 128k max output
  • $2 input / $10 output introductory through August 31, 2026, then $3 / $15
  • Balanced coding and tool-use tier
  • Best for: high-volume agents, chat surfaces, and cost-sensitive production where Opus 5 is overkill

What it does not win: top-end capability (Opus 5 and Fable 5 sit above it), and rock-bottom price (open-weight coders undercut it on output tokens).

GPT-5.6 Sol. Best for independently verified reasoning

OpenAI. Generally available July 9, 2026. Native vision and audio.

OpenAI’s flagship for reasoning-heavy work, and the July pick when you want an independently measured score. Sol posts 94.1% on GPQA Diamond per Artificial Analysis, an independent tracker, and holds the top Artificial Analysis Intelligence Index position among the GPT-5.6 tiers. Its price held at $5/$30 through the July 30 cuts that hit the cheaper tiers.

  • $5 input / $30 output per million tokens (unchanged July 30)
  • 94.1% GPQA Diamond (Artificial Analysis, independent)
  • Native vision and audio
  • Best for: reasoning-heavy applications, agentic work that wants a hosted flagship, broad ecosystem support

What it does not win: cost (Terra and Luna undercut it heavily after the July 30 cuts), and open-weight licensing (it is hosted only). OpenAI published no SWE-bench Verified figure at launch, so pair it with your own coding eval.

GPT-5.6 Terra and Luna. The price-cut mid and low tiers

OpenAI. Generally available July 9, 2026. Repriced July 30, 2026.

The two cheaper GPT-5.6 tiers, and where OpenAI spent its July pricing move. On July 30, Terra dropped 20% to $2/$12 and Luna dropped 80% to $0.20/$1.20, which makes Luna one of the cheapest hosted frontier-family tiers available. Both carry Artificial Analysis Intelligence Index positions below Sol (Terra 55, Luna 51 on the 100-point scale).

  • Terra: $2 input / $12 output per million tokens (cut 20% July 30 from $2.50/$15)
  • Luna: $0.20 input / $1.20 output per million tokens (cut 80% July 30 from $1/$6)
  • Both native vision and audio
  • Best for: high-volume production (Terra) and cost-sensitive bulk work (Luna) inside the OpenAI ecosystem

What it does not win: top-end reasoning (Sol leads the family), and the open-weight economics of the DeepSeek and Tencent tiers on the cheapest bulk work.

Gemini 3.6 Flash. Best cheap fast frontier

Google DeepMind. Generally available July 21, 2026.

The cheap, fast frontier option, and Google’s July speed release. Gemini 3.6 Flash runs at $1.50 input / $7.50 output per million tokens with a 1,048,576-token context window, and Google reports it using 17% fewer output tokens than 3.5 Flash for the same work, which lowers real cost below the sticker gap. Provider benchmarks put it at 58.7% on SWE-bench Pro and 83% on OSWorld Verified.

  • $1.50 input / $7.50 output per million tokens
  • 1,048,576-token context window, 65,536 max output
  • 17% fewer output tokens than 3.5 Flash (provider-reported)
  • Best for: high-volume, latency-sensitive agentic work that wants frontier-adjacent quality cheaply

What it does not win: top-end coding and reasoning (the Pro tiers and Opus 5 lead), and open-weight economics for the cheapest bulk output.

Grok 4.5. Best for reasoning with real-time data

xAI. Public July 8, 2026. 500K context. Proprietary, API-only.

xAI’s current flagship, and the pick when reasoning meets a need for current-events grounding. Grok 4.5 ships proprietary and API-only with a 500K-token context window (down from Grok 4.3’s 1M) at $2 input / $6 output, with prompt caching available. It pulls real-time data from X without a separate retrieval layer.

  • $2 input / $6 output per million tokens
  • 500K-token context window
  • Real-time data via X
  • Best for: research agents that need current data, reasoning-heavy workflows, X-native grounding

What it does not win: open-weight licensing (it is hosted only), and the longest context window (it dropped to 500K). Its coding and terminal benchmarks are single-source at launch, so verify them before you rely on them.

Top open-weight and Chinese-frontier LLMs in July 2026

The open-weight tier is where July set a record. Kimi K3 became the largest open-weight model ever released, DeepSeek and Tencent held the cost-performance floor, and Alibaba kept its flagship Max tier closed. Weigh license and real inference cost alongside the download size, because the biggest model is not the cheapest to run.

Kimi K3. The July open-weight record

Moonshot AI. API July 16, open weights July 27, 2026. Modified MIT (“Kimi K3 License”), 1M context.

The single most important open-weight release of July 2026, and the largest open-weight model ever shipped. Kimi K3 carries 2.8 trillion total parameters and a 1M-token context window under a Modified MIT license (Moonshot’s own “Kimi K3 License”), with a download near 1.56TB.

The license is permissive for most users but gates commercial use at scale: teams above $20M in annual model-as-a-service revenue or 100M monthly active users need a separate agreement with Moonshot and must display “Kimi K3” attribution.

Moonshot reports strong coding and reasoning numbers (GPQA Diamond 93.5%, SWE-bench Verified 76.8%, Terminal-Bench 2.1 88.3%), all provider-reported, so treat them as claims to reproduce.

The price reflects the size. Kimi K3 lists $3 input / $15 output per million tokens, which is expensive for an open model because a 2.8T-parameter network with a large active-parameter count is costly to serve. The value here is open weights at record scale, not the cheapest output.

  • Modified MIT license (the “Kimi K3 License”; large-scale commercial use over $20M revenue or 100M MAU needs a separate agreement), downloadable weights (~1.56TB)
  • 2.8T total parameters, 1M-token context window
  • Provider-reported GPQA Diamond 93.5%, SWE-bench Verified 76.8% (independent runs pending)
  • $3 input / $15 output per million tokens

Best for: teams that need open weights at frontier scale and can run or rent the inference, and who will benchmark the provider numbers on their own domain.

Skip if: you need an independently verified score today, or output price is your dominant constraint (DeepSeek and Tencent undercut it heavily).

DeepSeek V4-Pro and V4-Flash-0731. Best cost-performance on the open frontier

DeepSeek. V4-Pro generally available July 20; V4-Flash-0731 public beta July 31, 2026. MIT license.

The cost-performance anchors of the open tier, both MIT and both current in July. V4-Pro is the coder at 80.6% SWE-bench Verified (provider) and 91.6% LiveCodeBench (provider), a 1.6T mixture-of-experts with a 1M-token context window. V4-Flash-0731 is the routine-work beta at a fraction of the price.

  • V4-Pro: MIT, 1.6T / 49B active, 1M context; 80.6% SWE-bench Verified (provider); API around $0.27 / $1.10 (aggregator-reported, verify on the provider page)
  • V4-Flash-0731: MIT, 284B / 13B active, 1M context; $0.14 input / $0.28 output per million tokens
  • Both downloadable weights

Best for: high-volume coding and general work where output price dominates, and teams that can self-host or accept inference from DeepSeek’s API.

Skip if: you need an independent benchmark today (the published scores are provider-reported), or you want a single hosted flagship with vendor support.

Tencent Hunyuan Hy3. Cheapest permissive coder

Tencent. Generally available July 6, 2026. Apache-2.0, 256K context.

Tencent’s July open release, and the cheapest permissively licensed coder on this list. Hunyuan Hy3 is a 295B mixture-of-experts (192 experts, top-8 routing, 21B active) with a 256K-token context window, listed around $0.13 input / $0.53 output on OpenRouter. Its provider SWE-bench Verified of 74.4% is single-source, so name it and verify before you rely on it.

  • Apache-2.0 license, 256K-token context window
  • 295B total / 21B active mixture-of-experts
  • ~$0.13 input / $0.53 output per million tokens (aggregator-reported)
  • 74.4% SWE-bench Verified (provider, verify)

Best for: cost-sensitive coding on a truly permissive license, and teams that want an Apache-2.0 alternative to the MIT Chinese models.

Skip if: you need a published independent score, or a context window longer than 256K for whole-repository work.

GLM-5.2. The June headliner, still current

Z.ai / Zhipu. Generally available June 13, 2026. MIT license, 1M context. No July update.

June’s price-war headliner, carried into July unchanged. GLM-5.2 ships under an MIT license with a 1M-token context window and roughly 750B total parameters, and Z.ai did not ship a GLM-5.3 or GLM-5.5 in July, so this is the current Z.ai open flagship. Its coding claims are vendor-reported, as covered in the June edition.

  • MIT license, downloadable weights
  • ~750B total / ~40B active, 1M-token context window
  • No July version update

Best for: self-hosted or hosted coding agents that want an MIT license and a 1M context, with your own eval to confirm the vendor coding numbers.

Skip if: you need the newest open release (Kimi K3, DeepSeek V4-Pro, and Tencent Hy3 all shipped in July).

Qwen3.8-Max, Mistral Large 3, and Llama 4. The closed preview and the carryovers

Alibaba (preview July 19, 2026), Mistral AI (December 2, 2025), and Meta (April 2025).

Three names worth tracking, with different availability. Qwen3.8-Max previewed July 19 as an API-only model, so Alibaba’s flagship Max tier stays closed at launch while smaller Qwen sizes remain open; do not treat it as open-weight. Mistral Large 3 remains the strongest non-Chinese permissive option at Apache-2.0, 675B mixture-of-experts, and a 262K context window. Llama 4 Scout and Maverick carry unchanged from April 2025 under the Llama Community License.

  • Qwen3.8-Max: API-only preview, closed weights (not an open-weight pick)
  • Mistral Large 3: Apache-2.0, 675B / 41B active, 262K context
  • Llama 4 Scout / Maverick: 10M / 1M context, Llama Community License

Best for: a truly permissive non-Chinese license (Mistral Large 3), and long-context open deployments (Llama 4 Scout at 10M tokens).

For formal-math work specifically, Mistral also shipped Leanstral 1.5 on July 2 (Apache-2.0, 119B / 6B active, 256K context), a Lean 4 proof agent that posts 587/672 on PutnamBench. Treat it as a domain specialist, not a general LLM pick.

Top multimodal LLMs in July 2026

Multimodal is three separate races now: vision, image generation, and video generation, each with its own leaders and price floors. Two deprecations matter this month, so read the notes below before you build.

Vision (image and document understanding)

Use caseBest closedBest open
General image understandingGemini 3.1 ProQwen open VL family
Document / OCR / chart readingClaude Opus 5 / Fable 5Qwen open VL family
Long-context multi-imageGemini 3.6 Flash (1M tokens)Llama 4 Scout (10M)
Screen / UI understandingGPT-5.6 SolQwen open VL family

Every Western frontier model now handles vision natively, including Opus 5, Fable 5, GPT-5.6, Gemini 3.6 Flash, and Grok 4.5. The open column has closed the gap on the basics while still trailing on dense charts and technical diagrams, so confirm any open VL model on your own document distribution before you switch.

Image generation

The image-gen category stayed fragmented in July, and one flagship left the market. Pick by what the image is for.

Use caseBest pickWhy
General-purpose closedGPT Image 2Still the OpenAI image flagship
Google flagship imageGemini 3 Pro ImageTop Gemini image tier
Fast Google imageGemini 3.1 Flash ImageSpeed-tuned Gemini image model
Cheap image tier (new)Gemini 3.1 Flash-Lite ImageNew July 1, $0.034 per 1,000 images, ~4s
Open-weight / self-hostFLUX.2 ProOn your hardware, from Black Forest Labs
Alternative closedSeedream v5ByteDance image model

Two July notes. Google added Gemini 3.1 Flash-Lite Image (Nano Banana 2 Lite) on July 1 at $0.034 per 1,000 images. Black Forest Labs announced FLUX 3 on July 23, but only the video variant is in early access, so treat FLUX 3 image as announced, not shipped. Drop Imagen 4 and Imagen 4 Ultra entirely: they were deprecated June 15 with a hard shutdown on August 17, 2026.

Video generation

Video-gen leadership held through July, with one clear deprecation to plan around.

Use caseBest pickWhy
Best all-aroundGoogle Veo 3.1Leads prompt adherence, native audio, 4K output
Multi-shot storytellingKling 3.0Sequences with subject consistency across cuts
Native audio-video (preview)Gemini Omni FlashDeveloper API preview since June 30

For production video in July 2026, Veo 3.1 is the default for narrative scenes and Kling 3.0 for multi-shot work. ByteDance shipped Seedance 2.5 on July 31, but only on its consumer app, so the API is not live yet; hold it until the developer endpoint ships. Remove OpenAI Sora 2 from new builds: the web and app went dark April 26, and the API ends September 24, 2026.

Audio understanding

Native audio handling (speech as input, without a separate transcription step) is standard across the Western frontier models. Google’s Gemini Omni Flash adds native audio in developer preview, and GPT-5.6 and Gemini 3.1 Pro both handle speech-in directly, which preserves prosody and disambiguates accents that break transcribe-then-prompt pipelines. For the full speech stack, see the dedicated voice guide below.

Voice and audio (covered separately)

Voice AI has its own decision logic. STT (AssemblyAI, Deepgram, ElevenLabs, OpenAI), TTS (Cartesia, ElevenLabs, Deepgram, Hume), and voice-agent platforms (LiveKit, Pipecat, Vapi, Retell) follow different picks than text-only LLMs. The conversational latency budget alone, anchored on the ITU-T G.114 one-way delay recommendation, drives different choices than text-LLM workflows.

See Best Voice AI of July 2026 for the full STT, TTS, and voice-agent stack.

Top embeddings and retrieval models in July 2026

Embeddings have two production constraints: retrieval quality on your domain, and price per million tokens. The public leaderboards churn month to month, so the honest move is to name the durable options and run your own retrieval eval rather than chase a live board position.

Use caseDurable pickNotes
Contextualized retrieval (new)voyage-context-4Chunk embeddings with document context, $0.12 / 1M
Best general retrieval (closed)voyage-4-largeMixture-of-experts, Voyage’s top retrieval model
Multimodal retrievalgemini-embedding-001Text plus image, audio, and PDF in one model
OpenAI ecosystem defaulttext-embedding-3-largeStrong general default; -3-small for cheapest viable
Multilingual enterpriseCohere embed-v4Multilingual production coverage
Open-weight / self-hostQwen3-Embedding-8BStrong open multilingual option on your hardware

Voyage’s current pick is voyage-context-4 (June 29, contextualized chunk embeddings at $0.12 per million tokens) or voyage-4-large; voyage-3-large is superseded, so do not cite it as current.

We are deliberately not printing MTEB numbers here, because the board positions move often and several current figures did not clear a two-source check at publication. Shortlist two or three of the models above, score them on your own documents, and add a reranker for a few points of retrieval quality at small added cost.

Top coding-specific LLMs in July 2026

Coding is the highest-stakes category, and July pushed both ends: a cheaper capable pick at the top (Opus 5) and a new scale leader in open weights (Kimi K3). One caveat frames the whole table: independent SWE-bench Verified rankings were not confirmable in July, so every score below is provider-reported unless labeled otherwise.

Use caseTop pickScoreOutput $/M tokens
Balanced agentic codingClaude Opus 5Agentic-coding focus (no single provider SWE headline)$25
Raw capabilityClaude Fable 595.0% SWE-bench Verified (provider)$50
Open-weight scaleKimi K376.8% SWE-bench Verified (provider)$15
Cost-performance open coderDeepSeek V4-Pro80.6% SWE-bench Verified (provider)~$1.10
Cheapest Apache open coderTencent Hunyuan Hy374.4% SWE-bench Verified (provider)$0.53
Cheap fast frontier codingGemini 3.6 Flash58.7% SWE-bench Pro (provider)$7.50
Cheapest frontier-class openDeepSeek V4-Flash-0731Provider harness only$0.28

One label matters here. The independent SWE-bench Verified leaderboard did not return structured data at publication, and only Terminal-Bench 2.0 was fetchable, so we did not print an independent coding ranking for July. Every score above is provider-reported. The gap between a provider score and your reproduction is the number that decides your pick, so run a domain eval on your own repositories before you trust any single figure.

Top reasoning and math LLMs in July 2026

Reasoning is where July’s cleanest independent numbers live, because GPQA Diamond has an Artificial Analysis run that clears a two-source check. GPT-5.6 Sol leads at 94.1% GPQA Diamond (Artificial Analysis, independent), with Gemini 3.5 Flash at 92.2% and Claude Fable 5 at 92.6% on the same board.

On provider-reported reasoning, Gemini 3.1 Pro’s model card lists GPQA Diamond 94.3%, SWE-bench Verified 80.6%, and HLE 44.4% without tools, and Claude Opus 4.8’s system card lists GPQA Diamond 93.6%. Claude Opus 5’s GPQA reports span 84.1% to 93.7% depending on the effort setting, so we name the model and skip a single figure.

Meta’s closed Muse Spark posts HLE 58% in its highest-effort mode (provider). If reasoning is your core workload, weight the independent GPQA Diamond result above any single-source claim and run your own eval on a graduate-level set that matches your domain.

Best LLM for X: decision framework

Choose Claude Opus 5 if:

  • You want balanced frontier capability for agentic coding at half the Fable 5 price.
  • You need a 1M-token context window for long-horizon or whole-repository work.
  • You are on the Anthropic ecosystem and want the new production default.

Choose Claude Fable 5 if:

  • You need the highest available coding capability and cost is secondary.
  • Your work is hard multi-file reasoning where the capability ceiling pays for itself.

Choose GPT-5.6 Sol if:

  • Reasoning is your core workload and you want an independently measured score.
  • You want a hosted flagship with native vision, audio, and broad ecosystem support.

Choose Kimi K3 if:

  • You need open weights at frontier scale and can run or rent the inference.
  • A Modified MIT license and a 1M context matter more than the cheapest output price.

Choose DeepSeek V4-Pro or V4-Flash-0731 if:

  • Output price is the dominant constraint.
  • Your workload is primarily coding (V4-Pro) or high-volume routine work (V4-Flash-0731).
  • MIT licensing matters for downstream redistribution.

Choose Gemini 3.6 Flash if:

  • You want cheap, fast frontier-adjacent quality at $7.50 output.
  • A 1M-token context window and lower output-token counts cut your real cost.

Choose Grok 4.5 if:

  • Graduate-level reasoning and real-time data are both core.
  • You want X-native grounding without a separate retrieval layer.

Avoid Gemini 3.5 Pro, Qwen3.8-Max, FLUX 3 Image, and Seedance 2.5 for now. Gemini 3.5 Pro did not ship in July, Qwen3.8-Max is an API-only preview, FLUX 3’s image variant is announced but not released, and Seedance 2.5’s API is not live. None is a July production option.

Common mistakes when picking an LLM in July 2026

The four most expensive errors we see production teams make:

  1. Treating provider scores as independent. July’s SWE-bench Verified numbers for Kimi K3, DeepSeek V4-Pro, and Tencent Hy3 are provider-reported, and the independent leaderboard did not render. Confirm coding scores on your own repositories before you migrate.
  2. Assuming the biggest open model is the cheapest. Kimi K3 is the largest open-weight model ever, and at $15 output it costs more than several Western mid-tiers. Scale is not savings, so price the inference before you adopt it.
  3. Ignoring total cost of ownership. Listed price times token volume is the sticker number. Real cost includes retry rate, thinking and effort multipliers, and failure recovery, and a flaky agent triples the bill before you notice.
  4. Building on a preview as if it were generally available. Qwen3.8-Max, FLUX 3 Image, Seedance 2.5, and Gemini 3.5 Pro are not shipped production options in July. Build on generally available models and keep a fallback in any path that cannot afford downtime.

Recent platform updates

DateEventWhy it matters
July 1Claude Mythos 5 access restored for US orgs; Gemini 3.1 Flash-Lite ImageRestricted Fable-class access returns; new cheap image tier
July 2Mistral Leanstral 1.5 (Apache-2.0)Lean 4 formal-math specialist
July 6Tencent Hunyuan Hy3 (Apache-2.0)Cheapest permissive open coder
July 8xAI Grok 4.5 publicReasoning with real-time X data, 500K context
July 9OpenAI GPT-5.6 (Sol/Terra/Luna) generally availableFlagship family reaches GA
July 16Moonshot Kimi K3 APIThe record open-weight model goes live
July 19Alibaba Qwen3.8-Max preview (API-only)Flagship Max tier stays closed
July 20DeepSeek V4-Pro generally available (MIT)Cost-performance open coder
July 21Google Gemini 3.6 Flash + 3.5 Flash-Lite (no 3.5 Pro)Cheap fast frontier; 3.5 Pro slips
July 23FLUX 3 announced (early access, video only)Next FLUX generation, image not yet shipped
July 24Anthropic Claude Opus 5 generally availableNew $5/$25 mid-flagship at half Fable 5
July 27Kimi K3 open weightsLargest open-weight model ever released
July 30GPT-5.6 Terra/Luna price cutsTerra down 20%, Luna down 80%
July 31DeepSeek V4-Flash-0731 beta; Seedance 2.5 (consumer)Cheapest open beta; video model on consumer app only

How to actually pick one for production

The leaderboard is the wrong artifact to make a July 2026 production decision from. Three things to do instead:

  1. Run a domain reproduction. Take 100 to 500 of your actual production prompts, run them through your two or three candidate models with your harness, and score them with Future AGI evals or your own judge. For July specifically, this is how you separate Opus 5’s agentic-coding framing and Kimi K3’s provider scores from the number you will actually ship.
  2. Measure reliability under load. A public score is an aggregate over a fixed task set, and it does not predict variance across your prompts and repeated agent runs. Use Future AGI Simulate to stress-test agents across long sessions, where most frontier models lose a chunk of headline accuracy that benchmarks never surface.
  3. Cost-adjust every score. July’s spread runs from $0.28 to $50 per million output tokens. Compute score-per-dollar on your domain before you default to a brand, because Kimi K3’s record scale and DeepSeek’s cheap output only pay off if they hold up on your traffic.

July rewarded teams that treated model choice as a price, license, and reliability decision. The top of the market got cheaper with Opus 5 and the GPT-5.6 cuts, the open tier got larger with Kimi K3, and both moves reward the team that measures before it commits.

The practical answer is to shortlist by your dominant workload, then let your own reproduction and reliability numbers pick the winner. Put more effort into the eval loop above the model than into comparing leaderboard rows, because that is the layer that decides whether July’s cheapest pick or its most capable one actually works for you.

Sources

Frontier model launches (primary):

Open-weight and benchmarks:


Previous: Best LLMs of June 2026 ←. The open-weight coding price war went mainstream and Anthropic opened a new top tier.

Frequently Asked Questions

What is the best LLM in July 2026?

There is no single best LLM in July 2026. By use case: Claude Opus 5 is the balanced agentic-coding frontier at $5 input / $25 output per million tokens, half the price of Fable 5, with a 1M-token context window. Claude Fable 5 is the highest-capability pick at 95.0% SWE-bench Verified (provider-reported) on a $10/$50 tier. GPT-5.6 Sol leads independently measured reasoning at 94.1% GPQA Diamond (Artificial Analysis). Kimi K3 is the open-weight scale leader as the largest open-weight model ever shipped. Gemini 3.6 Flash is the cheap fast frontier at $1.50/$7.50. Pick by workload, not leaderboard rank.

What is the best open-source LLM in July 2026?

Kimi K3 (Moonshot, open weights July 27 2026, Modified MIT license, 2.8T total parameters, 1M context) is the July headline as the largest open-weight model ever released, at $3/$15 per million tokens. DeepSeek V4-Pro (MIT, 1.6T mixture-of-experts, 1M context, generally available July 20) is the cost-performance coder. DeepSeek V4-Flash-0731 (MIT) is the cheapest frontier-class option at $0.28 output. Tencent Hunyuan Hy3 (Apache-2.0, 295B, generally available July 6) is the cheapest permissive coder at about $0.53 output. Mistral Large 3 (Apache-2.0) is the strongest non-Chinese permissive option.

What is the best LLM for coding in July 2026?

Match the model to the workload. Claude Opus 5 is Anthropic's pick for complex agentic coding at $25 output, and Claude Fable 5 tops SWE-bench Verified at 95.0% (provider-reported) for the hardest work. DeepSeek V4-Pro reports 80.6% SWE-bench Verified (provider) at a fraction of the price, and Kimi K3 reports 76.8% (provider) with open weights. Independent SWE-bench Verified rankings were not confirmable in July, so treat every coding score as provider-reported and reproduce it on 100 to 500 of your own prompts before trusting it.

What new LLMs were released in July 2026?

July was a busy month. OpenAI completed the GPT-5.6 (Sol, Terra, Luna) rollout to general availability July 9, then cut Terra 20% and Luna 80% on July 30. Tencent shipped Hunyuan Hy3 July 6, xAI released Grok 4.5 public July 8, Moonshot shipped Kimi K3 (API July 16, open weights July 27), DeepSeek released V4-Pro July 20, Google shipped Gemini 3.6 Flash and 3.5 Flash-Lite July 21 (but not 3.5 Pro), and Anthropic launched Claude Opus 5 July 24. DeepSeek V4-Flash-0731 entered public beta July 31.

What is Claude Opus 5?

Claude Opus 5 (`claude-opus-5`, generally available July 24 2026) is Anthropic's new mid-flagship for complex agentic coding and enterprise work. It has a 1M-token context window and 128k max output, and is priced at $5 input / $25 output per million tokens, roughly half the cost of Claude Fable 5. Anthropic frames it around agentic-coding and enterprise benchmarks rather than a single headline score, so reproduce it on your own tasks before you budget around a specific number.

What is Kimi K3?

Kimi K3 (Moonshot AI) is the largest open-weight model ever released, generally available via API July 16 2026 and as open weights July 27 under a Modified MIT license. It has 2.8 trillion total parameters, a 1M-token context window, and a download near 1.56TB, priced at $3 input / $15 output per million tokens. Its published benchmark scores are provider-reported, so verify them on your own workload before adopting.

How much does Claude Opus 5 cost?

Claude Opus 5 costs $5 input / $25 output per million tokens, about half the price of Claude Fable 5 ($10/$50). It carries a 1M-token context window and 128k max output. That pricing places it below Fable 5 and above Claude Sonnet 5, whose introductory rate is $2/$10 through August 31 2026 then $3/$15.

Did OpenAI cut GPT-5.6 prices in July 2026?

Yes. On July 30 2026, OpenAI cut GPT-5.6 Terra 20% to $2 input / $12 output per million tokens (from $2.50/$15) and cut GPT-5.6 Luna 80% to $0.20 input / $1.20 output (from $1/$6). GPT-5.6 Sol, the flagship tier, held at $5/$30. The GPT-5.6 family reached general availability July 9 2026.
Related Articles
View all