Open-Source AI News: July 2026 LLM Release Tracker
A verified July 2026 open-source LLM release tracker separating announcements from downloadable weights, original models from derivatives, and license types.
Table of Contents
Last verified: 27 July 2026, 1:30 p.m. Pacific Time.
Nine notable LLM releases had downloadable weights in July by the time of this update. The list now includes Kimi K3: Moonshot announced it on 16 July, and the full checkpoint files became available on 27 July.
This tracker counts availability events, not headlines. A model qualifies when the weights are retrievable, the exact license is visible, and a primary source establishes what shipped.
The key July story is not model count. It is the spread from a 3B local agent model to Kimi K3 at 2.8T parameters, combined with three different license postures that headlines often collapse into “open source.”
For the terminology behind those labels, read open-source versus open-weight LLMs.
Which Open-Weight LLMs Shipped in July 2026?
The table separates original models from updates and derivatives. “Permissive open weight” means the checkpoint uses Apache-2.0; it does not claim that the full system has been assessed against OSAID 1.0.
| Model | Availability event | Release type | Parameters | License posture | Primary evidence |
|---|---|---|---|---|---|
| Leanstral 1.5 | 2 Jul | Family update | 119B / 6.5B active | Permissive open weight | Lab announcement and model card |
| Hy3 | 6 Jul | Original GA release | 295B / 21B active | Permissive open weight | Tencent announcement and model card |
| Bonsai 27B | 14 Jul | Low-bit derivative | 27B | Permissive open weight | Checkpoint and model card |
| Inkling | 15 Jul | Original model | 975B / 41B active | Permissive open weight | Lab announcement and model card |
| Nanbeige4.2-3B | 21 Jul | Original compact model | 3B non-embedding | Permissive open weight | Checkpoint and model card |
| Laguna S 2.1 | 21 Jul | Family update | 118B / 8B active | OpenMDW-1.1 | Lab announcement and license |
| Solar Open 2 | 22 Jul | Original model | 250B / 15B active | Custom Solar license | Lab announcement and license |
| KAT-Coder-V2.5-Dev | 23 Jul | Post-trained derivative | 35B / 3B active | Permissive open weight | Checkpoint and model card |
| Kimi K3 | 27 Jul weights | Original model | 2.8T / 104B active | Custom Kimi K3 license | Model card, files, and license |
The table is a release ledger, not a leaderboard. Parameter count does not compare intelligence, serving cost, or workload fit.
How Does the Release Ledger Avoid False “Launch” Dates?
Model releases have at least three timestamps:
- Announced: the lab publishes a post, demo, or benchmark claim.
- Weights available: the checkpoint files can be retrieved under visible terms.
- Serving supported: a documented runtime can load the exact architecture and precision.
Those events can happen on different days. Kimi K3 illustrates the lag: Moonshot announced it on 16 July, while the downloadable weight shards appeared on 27 July.
This tracker includes a model in the monthly release table when event two occurs. Announcement context belongs in the notes, while serving readiness gets a separate check.
We do not use Hugging Face downloads as an adoption threshold. The Hub explains that its download statistic counts qualifying file requests, not unique users or production deployments.

Which July Releases Changed the Frontier?
Three releases changed the shape of the month: Kimi K3 at the top end, Inkling as a multimodal base, and Hy3 as a more deployable large mixture of experts.

Kimi K3 Made July’s Largest Checkpoint Downloadable
Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model with 104B active parameters, native image input, and a one-million-token context window.
The checkpoint is real, not a countdown page: the repository file tree contains the weight shards and configuration needed for download.
Its license is not Apache-2.0. The Kimi K3 License allows broad use, modification, distribution, fine-tuning, and sale, then adds two scale conditions:
- a model-as-a-service business above $20 million in aggregate annual revenue needs a separate agreement for commercial use;
- a product above 100 million monthly active users or $20 million monthly revenue must display “Kimi K3” prominently.
That makes Kimi K3 open weight under custom terms, not OSAID-open by default.
The operational catch is size. “Downloadable” does not mean practical for one team to serve. Check the exact quantization, runtime support, memory topology, and context target before estimating infrastructure.
Inkling Is the Month’s Multimodal Open-Weight Base
Thinking Machines describes Inkling as a 975B-parameter model with 41B active parameters. It accepts text, images, and audio and exposes a controllable reasoning-effort trade-off.
The lab is refreshingly direct in its announcement: it does not claim Inkling is the strongest model overall. It positions the model as a balanced, adaptable base.
The checkpoint uses Apache-2.0. That is broad permission over the released artifact, not automatic proof that every training artifact satisfies OSAID.
Hy3 Put a Large General Model Under Apache-2.0
Tencent released Hy3 on 6 July with 295B total parameters, 21B active parameters, and a 256K context window.
The useful change from its preview is the license posture. The current Hy3 model card identifies Apache-2.0 and includes deployment instructions.
Hy3 is still a large serving job. Sparse activation reduces per-token compute relative to a dense 295B model, but it does not shrink the stored checkpoint to 21B parameters.
Which Specialist and Efficient Releases Matter?
The rest of the month is more useful when grouped by job than by parameter count.
Leanstral 1.5 Targets Formal Verification
Mistral’s Leanstral 1.5 is a 119B mixture-of-experts model with 6.5B active parameters, designed for Lean 4 proof engineering.
This is not a general coding recommendation. It matters for formal proofs, autoformalization, and code-verification workflows where a specialized model can beat a broader assistant.
Bonsai 27B Makes a Derivative the Product
Bonsai is not a new 27B foundation-model pretraining run. It is a low-bit derivative built on a Qwen3.6-27B backbone, packaged for local hardware.
That distinction belongs in a tracker. Derivatives can be valuable releases, but counting one as a new base model inflates the apparent number of original models.
The model card reports one-bit and ternary builds plus long-context memory measurements. Treat those as vendor-reported until reproduced on your hardware.
Nanbeige4.2-3B Pushes the Small End
Nanbeige4.2-3B is a compact agentic model built around a looped-transformer design that reuses its layer stack.
At this scale, the interesting question is not whether it beats Kimi K3. It is whether a local deployment can complete your tool-use and office tasks within a memory and latency budget.
The model card includes limitations and an Apache-2.0 checkpoint.
Laguna S 2.1 Tests OpenMDW in Practice
Laguna S 2.1 has 118B total and 8B active parameters, a one-million-token context window, and a coding focus.
Its OpenMDW-1.1 license grants broad rights over the model materials provided under it. It also terminates grants when a user voluntarily brings certain patent or copyright infringement claims.
OpenMDW does not promise that every training artifact exists in the release. The OpenMDW FAQ says the license applies to the materials a provider places under it.
Solar Open 2 Adds a Branding Rule
Solar Open 2 is a 250B model with 15B active parameters and a one-million-token context window.
The custom Upstage Solar License adopts Apache-like grants, then requires a distributed derivative model’s name to begin with “Solar” and related interfaces or documentation to display “Built with Solar.”
That is workable for many teams. It is still a condition your legal and product groups need to see before fine-tuning.
KAT-Coder-V2.5-Dev Is a Coding Derivative
KAT-Coder-V2.5-Dev is a 35B mixture-of-experts model with 3B active parameters and an Apache-2.0 checkpoint.
Its card states that this open-weight release contains only the language-model weights, not the vision components. Runtime flags may need to disable multimodal initialization.
That is the kind of deployment detail a release tracker should surface. A successful download is not the same as a successful server start.
How Should You Read the License Column?
Use three labels rather than one “open” badge.
Permissively Licensed Open Weight
The checkpoint uses an OSI-approved software license such as Apache-2.0. This gives broad rights over the checkpoint but does not establish that the full training system meets OSAID.
July examples: Leanstral 1.5, Hy3, Bonsai, Inkling, Nanbeige4.2-3B, and KAT-Coder-V2.5-Dev.
Model-Specific Open License
The license is designed for model materials and grants broad rights, but its scope and termination terms need direct review.
July example: Laguna S 2.1 under OpenMDW-1.1.
Custom-License Open Weight
The weights are downloadable, but the vendor adds scale, branding, acceptable-use, or redistribution conditions.
July examples: Solar Open 2 and Kimi K3.
None of these labels says whether the model is good. They tell procurement and engineering which question comes next.
What Should You Verify Before Serving a New Model?
Run five checks in order.
Verify the Artifact
Pin the model ID and revision. Save the license and model card you reviewed. Confirm that all expected shards and tokenizer files are present.
Verify Runtime Support
Check whether the exact architecture is supported by the current release of your serving engine. The vLLM guide explains the serving layer; the self-hosted vLLM evaluation guide covers the rollout checks.
Reproduce a Minimal Load
Load the intended precision on the intended hardware. Record GPU memory, startup time, time to first token, throughput, and failure logs.
Do this before estimating cost from active-parameter count.
Test Your Workload
Use a fixed dataset across every candidate. Measure task quality, grounding, structured output, tool calls, safety, latency, and cost.
Public leaderboards are discovery tools. The state of LLM benchmarking and benchmarks-versus-production-evals guide explain why the model-card winner can still lose on your traffic.
Stage the Rollout
Start offline, then shadow traffic, then route a small percentage with a rollback rule. Keep the license record next to the model revision.
FutureAGI’s Evaluate platform provides the UI-based dataset and eval-run surface for this comparison; the serving and license decisions remain yours.
What Does July 2026 Tell Us About Open Models?
July did not produce one clean “best open-source model.” It produced a more useful map:
- Kimi K3 expanded the downloadable frontier to 2.8T parameters under custom terms.
- Inkling offered a large multimodal base under Apache-2.0.
- Hy3 put a general 295B mixture-of-experts checkpoint under Apache-2.0.
- Nanbeige and Bonsai pushed local deployment from different directions.
- OpenMDW, Solar, and Kimi showed why one license column is not enough.
The June LLM roundup provides the previous-month context. For July, the right adoption question is not “which model had the biggest launch?”
It is: which downloadable artifact can we legally use, reliably serve, and prove on our own workload?
Frequently Asked Questions About Open-Source LLM Releases in July 2026
What are the latest open-source LLM releases in July 2026?
Was Kimi K3 released in July 2026?
Are Apache-2.0 model weights open-source AI?
What is the difference between a model announcement and a weight release?
How should a team evaluate a new open-weight LLM?
Open-source and open-weight LLMs are not the same. Use the OSI test to check model weights, code, data information, licenses, and deployment rights.
Best LLMs of June 2026 by use case: Claude Fable 5 for raw coding, GLM-5.2 for open-weight value, GPT-5.5 for agents, Gemini 3.1 Pro for long-context multimodal.
MMLU, GSM8K, SWE-bench Verified, BFCL, tau-bench, GPQA, ARC-AGI-2, Chatbot Arena. What each measures, where each breaks, triangulate-plus-private 2026.