AI Model & Benchmark Watch — July 24, 2026
Google finally shipped something — Gemini 3.6 Flash, 3.5 Flash-Lite, and a government-only "Flash Cyber" — but still no Gemini 3.5 Pro, while the White House escalated its Kimi K3 fight with a direct…
AI Model & Benchmark Watch — July 24, 2026
Google finally shipped something — Gemini 3.6 Flash, 3.5 Flash-Lite, and a government-only "Flash Cyber" — but still no Gemini 3.5 Pro, while the White House escalated its Kimi K3 fight with a direct accusation that Moonshot AI distilled Anthropic's Fable model using smuggled Nvidia chips.
Overview
The frontier itself barely moved this week — the Artificial Analysis Intelligence Index top 3 (Claude Fable 5, GPT-5.6 Sol, Kimi K3) is unchanged from July 17 — but the story around it got a lot louder. Google used its Gemini 3.6 Flash launch (July 21) to paper over the still-missing Gemini 3.5 Pro, quietly teasing "Gemini 4" instead of naming a Pro ship date. Alibaba previewed a 2.4-trillion-parameter Qwen3.8-Max (July 19) that it claims trails only Claude Fable 5, though it published zero benchmark numbers to back that up. And the geopolitical temperature around open-weight Chinese models spiked sharply: White House OSTP director Michael Kratsios accused Moonshot AI on July 22 of covertly distilling Anthropic's Fable model to train Kimi K3 and of routing around export controls to access banned Nvidia GB300 chips via Thailand — with Treasury now threatening sanctions. Meanwhile the DeepSeek V4 migration deadline flagged in last week's edition arrived on schedule: legacy deepseek-chat/deepseek-reasoner endpoints retire today, July 24, at 15:59 UTC.
New & Updated Models (July 17–24)
Gemini 3.6 Flash, 3.5 Flash-Lite & 3.5 Flash Cyber — Google DeepMind (July 21, 2026)
Who / License: Google DeepMind; closed.
What's notable: Google's answer to a third straight missed Gemini 3.5 Pro deadline was a Flash-tier refresh instead — no Pro ship date was given, and Google instead teased an upcoming "Gemini 4." Gemini 3.6 Flash is pitched as the workhorse tier: 1M-token context, gains in coding/knowledge/multimodal work, and up to 17% lower token usage than 3.5 Flash. It scores 50 on the Artificial Analysis Intelligence Index — well above the ~31 category average — with 90.4% on GPQA Diamond and 78% on SWE-bench Verified (58.7% on SWE-bench Pro). 3.5 Flash-Lite is the new budget option in the same family. 3.5 Flash Cyber is a specialized model fine-tuned for finding and fixing security vulnerabilities, restricted to governments and trusted partners in a limited pilot — not publicly available.
Pricing: Gemini 3.6 Flash — $1.50 / M input, $7.50 / M output (down from 3.5 Flash's $9.00 output rate). Context: 1M tokens.
Source: Google — Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber · TechCrunch — Google releases three new Gemini models, but no 3.5 Pro · 9to5Google — Gemini 3.6 Flash launch, teases Gemini 4 · Artificial Analysis — Gemini 3.6 Flash · officechai — Gemini 3.6 Flash benchmarks
Qwen3.8-Max Preview — Alibaba (July 19, 2026)
Who / License: Alibaba; closed preview endpoint (open weights "promised soon," no date or license named).
What's notable: Alibaba's first multimodal model above 1 trillion parameters — 2.4T total, sparse MoE, processing text, image, video, and documents. Previewed at WAIC in Shanghai, days after Kimi K3's launch. Alibaba claims it trails only Claude Fable 5 in overall capability, but as of this writing Alibaba has published no benchmark numbers, technical report, Artificial Analysis entry, or pricing for Qwen3.8 itself — only the vendor claim. Treat this as an unverified preview, not a leaderboard entry, until real scores land. Accessible today only via Alibaba's Token Plan, Qoder, and QoderWork platforms as "Qwen3.8-Max-Preview."
Source: Bloomberg — Alibaba's Qwen unveils preview of flagship AI model, shares rise · MarkTechPost — Alibaba previews Qwen3.8-Max, 2.4T-parameter multimodal model · South China Morning Post — Alibaba says newest Qwen model is second only to Fable 5 · techsy.io — Qwen3.8: 2.4T parameters, open weights, no benchmarks
Kimi K3 open weights — still pending (due July 27)
Who / License: Moonshot AI; API-only proprietary today, Modified MIT weights promised July 27.
What happened: No change from last week's edition — Kimi K3's 2.8T-parameter weights have not shipped yet; the July 27 date (3 days out) still stands per Moonshot's own tech blog. What's new is the controversy around it: see Benchmark & Leaderboard Movement below for the White House distillation accusation.
Source: Kimi K3 Tech Blog — Open Frontier Intelligence
OpenAI Presence — OpenAI (July 22, 2026)
Who / License: OpenAI; closed enterprise product (not a new base model).
What's notable: An enterprise agent-deployment product for scoped, policy-governed voice and chat agents in customer support, outbound sales, and internal IT — not a model release, but notable as OpenAI's move into the "boots-on-the-ground" agent-deployment/consulting business. OpenAI says Presence powers its own English-language phone support line, resolving 75% of inbound issues without human escalation. Built on GPT-5.6 and Codex; uses a Codex-based tool to propose workflow updates from production data.
Source: OpenAI — Introducing OpenAI Presence · The Register — OpenAI tries the consulting path with Presence · StreetInsider — OpenAI launches Presence
Head-to-Head — Current Frontier
The table below reflects the state of play as of July 23–24, 2026. AA Index = Artificial Analysis Intelligence Index (BenchLM.ai snapshot, data verified July 23). Arena Elo = Arena.ai (formerly LMArena) text leaderboard; models in the pool fewer than ~2 weeks have low vote counts and scores will keep shifting — marked with †; some figures are carried from the June/early-July snapshot where no fresher data was found (marked ‡). SWE-bench Pro column throughout (not comparable to Verified scores — see footnotes for models where only Verified is published). "—" = not publicly confirmed or not yet evaluated.
| Model | Org | Open? | Arena Elo (Text) | AA Index | GPQA Dia | SWE-bench Pro | Terminal-Bench 2.1 | Context | $ / M in/out |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | No | ~1,525‡ | 59.9 (#1) | — | 80.3% | 86.0% | 1M | $10 / $50 |
| GPT-5.6 Sol | OpenAI | No | ~1,465 †‡ | 58.9 (#2) | 94.1% | 64.6% | 88.8% / 91.9%§ | 1M+ | $5 / $30 |
| Kimi K3 | Moonshot AI | Yes* | 1,486 † | 57.1 (#3) | 93.5% | — (67.5% DeepSWE) | 88.3% | 1M | $3 / $15 |
| Claude Opus 4.8 | Anthropic | No | ~1,510‡ | 55.7 (#4) | — | 69.2% | ~85.0% | 1M | $5 / $25 |
| GPT-5.6 Terra | OpenAI | No | — | 55.0 (#5) | — | 63.4% | 87.4% | 1M+ | $2.50 / $15 |
| Grok 4.5 | xAI | No | ~1,496‡ | 53.8 (#7) | 93.0% | 64.7% | — | 500K | $2 / $6 |
| Gemini 3.6 Flash | No | — | 50.0 | 90.4% | 58.7% (78% Verified) | — | 1M | $1.50 / $7.50 | |
| GLM-5.2 | Zhipu / Z.ai | Yes (MIT) | — | 51.1 | 91.2% | 62.1% | 82.7% | 1M | $1.40 / $4.40 |
| DeepSeek V4-Pro | DeepSeek | Yes (MIT) | ~1,410‡ | 44.0 | — | — (80.6% Verified) | — | 1M | $0.44 / $0.87 |
| Claude Sonnet 5 | Anthropic | No | — | 53.4 (#9) | — | 63.2% | 80.4% | 1M | $2 / $10¶ |
*Kimi K3: API available; weights still pending as of July 24 — expected July 27 under Modified MIT. †Arena Elo: models that entered the pool within the last ~2–3 weeks have too few votes for stable rankings. ‡Arena Elo: no update found this week beyond the July 17 edition's figures; treat as directional, not current-day exact. §GPT-5.6 Sol (standard): 88.8%; Sol Ultra (4-agent Codex mode): 91.9%. ¶Claude Sonnet 5 at introductory pricing through August 31, 2026; then $3 / $15.
Reading the table: The top of the board is frozen — Fable 5, Sol, and K3 hold the same AA Index #1–#3 they held last week, and no new model challenged that order. The one real newcomer, Gemini 3.6 Flash, lands respectably (AA Index 50, above the ~31 category average for its tier) but is a mid-tier Flash model, not a Pro-class challenger — Google still has no answer at the frontier while it waits on Gemini 3.5 Pro. On pure GPQA Diamond, Grok 4.5's 93.0% edges past Kimi K3 (93.5% — statistically close) into the crowded 90%+ tier alongside Sol (94.1%) and GLM-5.2 (91.2%); the benchmark remains near-saturated at the top. DeepSeek V4-Pro's AA Index (44.0, freshly sourced this week) is notably lower than GLM-5.2's (51.1) despite similar GPQA figures reported elsewhere — a reminder that the AA Index is a composite across agentic, coding, and scientific tasks, not just knowledge recall, and DeepSeek's agentic/coding sub-scores are weaker. Fable 5's SWE-bench Pro lead (80.3%, ~11 points over Opus 4.8) remains the single clearest argument for its $10/$50 premium.
Benchmark & Leaderboard Movement
- White House accuses Moonshot AI of distilling Anthropic's Fable to build Kimi K3 (July 22): OSTP director Michael Kratsios said the administration has evidence Moonshot ran a covert, industrial-scale distillation platform against US models — including Anthropic's Fable — while rotating access methods to avoid detection, and separately accessed banned Nvidia GB300 chips via Thailand. Treasury is now threatening sanctions. This is the first time a senior US official has named a specific Chinese lab and a specific American model in a distillation accusation, and it directly implicates the model (Kimi K3) that displaced Claude Fable 5 atop Arena's Frontend Code leaderboard just six days earlier. Kimi K3's benchmark standing hasn't moved, but its provenance is now a live diplomatic dispute — worth watching ahead of the July 27 open-weights release.
- AA Intelligence Index top 3 holds steady: Claude Fable 5 (59.9), GPT-5.6 Sol (58.9), and Kimi K3 (57.1) are unchanged from the July 17 snapshot to the July 23 data pull — the first quiet week at the top since the index's post-June-launch churn.
- Gemini 3.6 Flash enters mid-pack: At AA Index 50.0, Gemini 3.6 Flash slots just behind GLM-5.2 (51.1) and ahead of GPT-5.6 Luna (51.2 — essentially tied) — a Flash-tier model now benchmarking in open-weight-frontier territory, though still well off the top-3 cluster.
- Qwen3.8-Max Preview: claim without evidence. Alibaba's "second only to Fable 5" claim has zero published benchmark backing as of July 23 — no technical report, no Artificial Analysis listing, no third-party eval. This is the loudest unverified claim of the week; treat it as marketing until real numbers land.
- DeepSeek V4 migration deadline hits today: The
deepseek-chat/deepseek-reasonerlegacy endpoint names — flagged as retiring in last week's edition — go dead today, July 24, at 15:59 UTC, with no grace period. Both names have been silently routed todeepseek-v4-flashduring the grace window; anything still calling the old names needs to be ondeepseek-v4-proordeepseek-v4-flashbefore the cutoff. - Grok 4.5 third-party benchmarks firm up: Independent coverage this week puts Grok 4.5 at 93.0% GPQA and 64.7% SWE-bench Pro — both now competitive with GPT-5.6 Sol's figures (94.1% / 64.6%), closing the gap from the sparse, mostly-self-reported picture in the July 8 edition.
Analysis
For agentic coding, nothing dislodged Claude Fable 5 this week — its 80.3% SWE-bench Pro score is still the clearest reason to pay $10/$50. Grok 4.5's newly-firmed 64.7% now edges out GPT-5.6 Sol (64.6%) in that specific metric, which is worth noting for teams already on xAI's stack, though the two are functionally tied. For reasoning/knowledge work, the 90–94% GPQA Diamond tier is crowded and close (Sol 94.1%, Grok 4.5 93.0%, Kimi K3 93.5%, GLM-5.2 91.2%, Gemini 3.6 Flash 90.4%) — pick on price and context rather than raw score. For cheap-and-fast production, Gemini 3.6 Flash ($1.50/$7.50) is a genuinely new, well-priced mid-tier option that beats its predecessor on cost while landing a competitive AA Index score; DeepSeek V4-Pro ($0.44/$0.87, MIT) remains the extreme-budget, self-hostable floor.
The open-vs-closed gap is unchanged from last week: Kimi K3 (open-pending) still sits between Sol and Opus 4.8 on the AA Index, and GLM-5.2 remains the most mature fully-open option in daily use. But this week is a reminder that "open" is getting geopolitically loaded — the White House's distillation accusation against Moonshot, arriving days before K3's promised weight release, means builders evaluating K3 for self-hosting should watch the July 27 date for both the weights themselves and any US regulatory response.
Sources
- Google — Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
- TechCrunch — Google releases three new Gemini models — but no 3.5 Pro
- 9to5Google — Google launches Gemini 3.6 Flash and 3.5 Flash-Lite, teases Gemini 4
- droid-life — Google drops Gemini Flash 3.6, teases huge Gemini 4 release
- Artificial Analysis — Gemini 3.6 Flash intelligence, performance & price
- officechai — Gemini 3.6 Flash beats Gemini 3.5 Flash, Gemini 3.1 Pro on most benchmarks
- BenchLM.ai — Gemini 3.6 Flash benchmarks, pricing & speed
- Bloomberg — Alibaba's Qwen unveils preview of flagship AI model, shares rise
- MarkTechPost — Alibaba previews Qwen3.8-Max, a 2.4T-parameter multimodal model
- South China Morning Post — Alibaba says newest Qwen AI model is second only to Fable 5
- techsy.io — Qwen3.8: 2.4T parameters, open weights, no benchmarks
- MLQ News — Alibaba launches Qwen 3.8 with 2.4 trillion parameters, claims near-frontier performance
- Kimi K3 Tech Blog — Open Frontier Intelligence
- TechCrunch — Treasury threatens sanctions after White House claims Moonshot distilled Anthropic's Fable
- Cryptopolitan — White House accuses Moonshot of distilling Anthropic's Fable for Kimi K3
- CyberScoop — White House accuses Chinese company of distilling Anthropic's Fable
- The Hill — White House official accuses Chinese startup of distilling Anthropic model, accessing banned Nvidia chips
- TFTC — White House: Moonshot AI accessed banned Nvidia GB300 chips to build Kimi K3
- OpenAI — Introducing OpenAI Presence
- The Register — OpenAI tries the consulting path with Presence
- StreetInsider — OpenAI launches Presence, an enterprise AI agent deployment product
- DeepSeek API Docs — V4 Preview Release and endpoint retirement notice
- Enterprise DNA — DeepSeek API: migrate before July 24 or integrations break
- TECHi — DeepSeek's legacy model names retire today, and the easy fix has a catch
- Digital Applied — DeepSeek's July 24 API cutoff: the alias migration playbook
- MarkTechPost — Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: open trillion-scale MoE models compared
- BenchLM.ai — Artificial Analysis Intelligence Index leaderboard (July 23 data)
- Artificial Analysis — Intelligence Index leaderboard
- llm-stats.com — GLM-5.2 vs Grok 4.5 benchmarks, pricing
- Blockchain.News — Kimi K3 tops Arena, but limits matter
- thedeveloperstory — Kimi K3 tops Arena's Frontend Coding leaderboard
- Local AI Master — LMArena leaderboard 2026 (live snapshot)
- Build Fast with AI — AI News Today July 23, 2026: 16 biggest stories
- TechCrunch — Anthropic updates Claude voice mode with more capable models
- decodethefuture.org — Claude Mythos 5: current status, access and Fable 5
- llm-stats.com — AI model leaderboard and updates, July 2026
More from News