July 2026: The 19 Days When Every Major AI Lab Shipped — a Timeline of What Actually Happened

July 2026 packed an unusual amount of AI model news into nineteen days. Between July 9 and July 27, OpenAI, Moonshot AI, Google DeepMind, and Anthropic each shipped models at or near the top of their lineups. By the loose rhythm of one major release per company every few months, this was clearly different.
The problem with a month like that is simple: you can follow each headline and still lose the plot. At least eight new model names appeared in this window, spanning different price bands and different roles inside the same company. This article threads the whole month onto a single timeline. Every date and number below comes from primary sources or primary reporting, and wherever a performance claim appears, we say whose claim it is.
- The nineteen days, in order
- OpenAI: one generation, three tiers
- Moonshot: an announcement, then the weights
- Google: the day Flash was the whole story
- Anthropic: same price, higher ceiling
- One neutral yardstick
- Trend one: you now choose how hard the model thinks
- Trend two: three opposite answers on cyber capability
- Regulation moved availability, too
- What this means at the workbench
- Three questions the month left behind
The nineteen days, in order
| Date | What happened |
|---|---|
| June 26 | OpenAI opens GPT-5.6 as a limited preview for trusted partners |
| July 9 | GPT-5.6 reaches general availability in three tiers: Luna, Terra, Sol |
| July 16 | Moonshot AI announces Kimi K3 |
| July 21 | Google DeepMind releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber |
| July 24 | Anthropic releases Claude Opus 5 |
| July 27 | Moonshot AI publishes the Kimi K3 weights |
| July 30 | OpenAI cuts GPT-5.6 Luna and Terra prices (Luna by 80 percent) |
Fifteen days separate the GPT-5.6 general release from Claude Opus 5, with Kimi K3 and three Gemini models landing in between. No company can re-plan a frontier release on that notice, so the clustering is better read as development cycles converging than as a reaction chain. Worth noting as background: Anthropic had already shipped Claude Opus 4.8 on May 28 and Claude Fable 5 on June 9, making Opus 5 the last step of a rapid sequence. That run-up is covered in our Claude summer catch-up.
OpenAI: one generation, three tiers
On July 9, OpenAI made GPT-5.6 generally available as three tiers named Luna, Terra, and Sol, in ascending order of capability. Launch pricing per million tokens was 1 dollar in and 6 dollars out for Luna, 2.50 and 15 for Terra, and 5 and 30 for Sol. Then on July 30 OpenAI cut prices: Luna now costs 0.20 in and 1.20 out, Terra 2 and 12, with Sol unchanged. The input spread between bottom and top tier went from five times to twenty-five times.
OpenAI calls Sol its best coding model to date and also its most capable cybersecurity model, citing a 54 percent gain in token efficiency for coding work. All of these are vendor claims rather than independent verification, so read them as a statement of where OpenAI sees its main battleground. Terra is positioned as competitive with the previous GPT-5.5 at a lower price. For how the lineup looked before this shift, see our ChatGPT guide from May.
Moonshot: an announcement, then the weights
Moonshot AI announced Kimi K3 on July 16 and published the weights on Hugging Face eleven days later, on July 27 — two separate dates that are easy to conflate. The headline numbers: 2.8 trillion total parameters in a mixture-of-experts design with 104 billion active per token, a context window of 1,048,576 tokens, and a download of about 1.56 terabytes. API pricing is 3 dollars in and 15 out per million tokens, with cached input at 0.30.
The license deserves care. This is an open-weight release, not open source. Hugging Face lists a custom Kimi K3 License rather than a standard identifier. Individuals, startups, and most companies can use, modify, fine-tune, and redistribute freely, but two revenue-linked conditions apply: model-as-a-service businesses earning over 20 million dollars in twelve months from hosting it need a separate agreement with Moonshot, and products with over 100 million monthly active users or 20 million dollars in monthly revenue must display the Kimi K3 name. A reasoning_effort parameter offers low, high, and max, with max as the default. At 1.56 terabytes, running it at home is not realistic — our piece on running open models locally gives a sense of what personal hardware actually handles.
Google: the day Flash was the whole story
On July 21, Google DeepMind released three models at once: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. Google says the new Flash improves coding, knowledge work, and multimodal performance while cutting output token usage by up to 17 percent — again, a vendor figure. Flash-Lite is pitched as the cheapest, fastest option in its class. Flash Cyber is different in kind: a model fine-tuned to find and fix security vulnerabilities, offered only to governments and trusted partners as a limited-access pilot.
What drew the most discussion was what did not ship: no update to the flagship Gemini Pro, last refreshed in February 2026 — five months earlier. Bloomberg reported on July 16 that the 3.5 Pro launch had slipped after missing internal performance goals. On release day, Google DeepMind product lead Logan Kilpatrick said 3.5 Pro was being tested with partners and should land soon. The previous flagship is covered in our Gemini 3.1 Pro guide.
Anthropic: same price, higher ceiling
On July 24, Anthropic released Claude Opus 5 at 5 dollars in and 25 out per million tokens — unchanged from Opus 4.8 and exactly half of Claude Fable 5 at 10 and 50. The pitch: a half-price model that matches the flagship on many tasks. Specs include a 1-million-token context window as both default and maximum, and a 128,000-token output cap. One behavioral reversal matters for migration: thinking is now on by default, where Opus 4.8 ran without it unless asked. Since the output cap covers thinking plus answer, tight limits tuned for no-thinking setups can silently truncate responses.
Anthropic also says its safety classifiers fire about 85 percent less often than on Fable 5, and introduced Automatic Fallbacks, a beta feature that routes a flagged request to a less restricted model instead of returning an error. Fewer refusals, plus a landing pad when they happen — a design direction we discussed in our piece on fallback design.
One neutral yardstick
Vendor benchmarks cannot be compared across vendors, so here is one independent cross-cutting measure: the Artificial Analysis Intelligence Index, as of early August 2026, at max effort settings.
| Model | Score |
|---|---|
| Claude Opus 5 (max) | 63 |
| Claude Fable 5 (max) | 62 |
| GPT-5.6 Sol (max) | 61 |
| Kimi K3 | 60 |
| Claude Opus 4.8 (max) | 57 |
Three things stand out. First, the top four sit within three points of each other — a pack, not a runaway leader. Second, Claude Opus 5 edges out the twice-as-expensive Claude Fable 5, which backs Anthropic’s half-price claim at least on this measure. Third, the open-weight Kimi K3 at 60 beats the commercial Claude Opus 4.8 at 57. Rankings move week to week — Kimi K3 was reported third at release and slipped to fourth after Opus 5 landed — so treat this as a snapshot, not a verdict.
Trend one: you now choose how hard the model thinks
Put the four launches side by side and a shared design emerges: the user, not the vendor, decides how much computation each request gets. Claude Opus 5 exposes an effort setting with five levels from low to max. Kimi K3 has reasoning_effort with three. OpenAI expresses the same idea as tiers — Luna, Terra, Sol — with depth settings inside the upper ones. Google does it with the Flash and Flash-Lite split. The upside is real: routine formatting no longer has to pay deep-reasoning prices. The catch is that both failure modes are silent — think too hard and the bill grows, too little and quality drops, and neither raises an error. Our effort guide for Claude Opus 4.8 covers the practical split.
Trend two: three opposite answers on cyber capability
The same month produced a genuine split on security capabilities. OpenAI leads with them, billing Sol as its strongest cybersecurity model and selling it openly. Google built a dedicated model but restricted it to governments and vetted partners. Anthropic runs the other way, using safety classifiers that refuse in that territory — while making them fire far less often on Opus 5. Sell it, gate it, or refuse it: three answers inside nineteen days. The practical takeaway is not about which is right, but that the same request may pass on one model and be refused on another. If you build on these systems, decide in advance what happens when a request is declined.
Regulation moved availability, too
GPT-5.6 spent two weeks in a trusted-partner preview before its July 9 general release. TechCrunch reported that the US administration had asked for rollout restrictions in June over misuse concerns, and the gated availability of Gemini 3.5 Flash Cyber fits the same pattern. The lesson for practitioners is blunt: whether you can use a model is not decided by benchmarks alone. Coupling your workflow tightly to a single vendor means someone else’s policy becomes your outage.
What this means at the workbench
For makers doing 3D printing and design work, the honest answer is: do not rush to switch tools. With the top four packed between 63 and 60 on an independent index, no swap will feel different for everyday CAD help, and migrating prompts has a real cost. What does change things: effort settings, where dropping light requests to a low setting visibly cuts the bill while geometry checks deserve depth — see our hands-on CAD sessions for where that line falls. Million-token contexts are now standard on Claude Opus 5 and Kimi K3, which makes whole-project consultations practical: design notes, failure logs, filament data sheets, and slicer history in one pass. And with GPT-5.6 Luna at 0.20 in and 1.20 out after the July 30 cut, bulk chores like reading dimensions off photos or triaging print logs get very cheap. The full workflow map lives in our AI 3D design roadmap.
Three questions the month left behind
First: which tier, for which task? All four vendors handed users the dial, so pick your three most frequent tasks and assign each a depth. Second: what happens when you get refused? Different vendors now draw the line in different places, so a fallback path belongs in any production setup. Third: what do you optimize besides scores? Regulation and distribution policy moved availability this month, which argues for judging supply stability alongside benchmarks. The numbers in this article will drift; the habit of designing your tool usage, rather than chasing rankings, is the part that lasts.





