← July 2026
News 2026-07-24

AI Model & Benchmark Watch — July 24, 2026

Google finally shipped something — Gemini 3.6 Flash, 3.5 Flash-Lite, and a government-only "Flash Cyber" — but still no Gemini 3.5 Pro, while the White House escalated its Kimi K3 fight with a direct…

AI Model & Benchmark Watch — July 24, 2026

AI Model & Benchmark Watch — July 24, 2026

Google finally shipped something — Gemini 3.6 Flash, 3.5 Flash-Lite, and a government-only "Flash Cyber" — but still no Gemini 3.5 Pro, while the White House escalated its Kimi K3 fight with a direct accusation that Moonshot AI distilled Anthropic's Fable model using smuggled Nvidia chips.

Overview

The frontier itself barely moved this week — the Artificial Analysis Intelligence Index top 3 (Claude Fable 5, GPT-5.6 Sol, Kimi K3) is unchanged from July 17 — but the story around it got a lot louder. Google used its Gemini 3.6 Flash launch (July 21) to paper over the still-missing Gemini 3.5 Pro, quietly teasing "Gemini 4" instead of naming a Pro ship date. Alibaba previewed a 2.4-trillion-parameter Qwen3.8-Max (July 19) that it claims trails only Claude Fable 5, though it published zero benchmark numbers to back that up. And the geopolitical temperature around open-weight Chinese models spiked sharply: White House OSTP director Michael Kratsios accused Moonshot AI on July 22 of covertly distilling Anthropic's Fable model to train Kimi K3 and of routing around export controls to access banned Nvidia GB300 chips via Thailand — with Treasury now threatening sanctions. Meanwhile the DeepSeek V4 migration deadline flagged in last week's edition arrived on schedule: legacy deepseek-chat/deepseek-reasoner endpoints retire today, July 24, at 15:59 UTC.

New & Updated Models (July 17–24)

Gemini 3.6 Flash, 3.5 Flash-Lite & 3.5 Flash Cyber — Google DeepMind (July 21, 2026)

Who / License: Google DeepMind; closed.

What's notable: Google's answer to a third straight missed Gemini 3.5 Pro deadline was a Flash-tier refresh instead — no Pro ship date was given, and Google instead teased an upcoming "Gemini 4." Gemini 3.6 Flash is pitched as the workhorse tier: 1M-token context, gains in coding/knowledge/multimodal work, and up to 17% lower token usage than 3.5 Flash. It scores 50 on the Artificial Analysis Intelligence Index — well above the ~31 category average — with 90.4% on GPQA Diamond and 78% on SWE-bench Verified (58.7% on SWE-bench Pro). 3.5 Flash-Lite is the new budget option in the same family. 3.5 Flash Cyber is a specialized model fine-tuned for finding and fixing security vulnerabilities, restricted to governments and trusted partners in a limited pilot — not publicly available.

Pricing: Gemini 3.6 Flash — $1.50 / M input, $7.50 / M output (down from 3.5 Flash's $9.00 output rate). Context: 1M tokens.

Source: Google — Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber · TechCrunch — Google releases three new Gemini models, but no 3.5 Pro · 9to5Google — Gemini 3.6 Flash launch, teases Gemini 4 · Artificial Analysis — Gemini 3.6 Flash · officechai — Gemini 3.6 Flash benchmarks


Qwen3.8-Max Preview — Alibaba (July 19, 2026)

Who / License: Alibaba; closed preview endpoint (open weights "promised soon," no date or license named).

What's notable: Alibaba's first multimodal model above 1 trillion parameters — 2.4T total, sparse MoE, processing text, image, video, and documents. Previewed at WAIC in Shanghai, days after Kimi K3's launch. Alibaba claims it trails only Claude Fable 5 in overall capability, but as of this writing Alibaba has published no benchmark numbers, technical report, Artificial Analysis entry, or pricing for Qwen3.8 itself — only the vendor claim. Treat this as an unverified preview, not a leaderboard entry, until real scores land. Accessible today only via Alibaba's Token Plan, Qoder, and QoderWork platforms as "Qwen3.8-Max-Preview."

Source: Bloomberg — Alibaba's Qwen unveils preview of flagship AI model, shares rise · MarkTechPost — Alibaba previews Qwen3.8-Max, 2.4T-parameter multimodal model · South China Morning Post — Alibaba says newest Qwen model is second only to Fable 5 · techsy.io — Qwen3.8: 2.4T parameters, open weights, no benchmarks


Kimi K3 open weights — still pending (due July 27)

Who / License: Moonshot AI; API-only proprietary today, Modified MIT weights promised July 27.

What happened: No change from last week's edition — Kimi K3's 2.8T-parameter weights have not shipped yet; the July 27 date (3 days out) still stands per Moonshot's own tech blog. What's new is the controversy around it: see Benchmark & Leaderboard Movement below for the White House distillation accusation.

Source: Kimi K3 Tech Blog — Open Frontier Intelligence


OpenAI Presence — OpenAI (July 22, 2026)

Who / License: OpenAI; closed enterprise product (not a new base model).

What's notable: An enterprise agent-deployment product for scoped, policy-governed voice and chat agents in customer support, outbound sales, and internal IT — not a model release, but notable as OpenAI's move into the "boots-on-the-ground" agent-deployment/consulting business. OpenAI says Presence powers its own English-language phone support line, resolving 75% of inbound issues without human escalation. Built on GPT-5.6 and Codex; uses a Codex-based tool to propose workflow updates from production data.

Source: OpenAI — Introducing OpenAI Presence · The Register — OpenAI tries the consulting path with Presence · StreetInsider — OpenAI launches Presence


Head-to-Head — Current Frontier

The table below reflects the state of play as of July 23–24, 2026. AA Index = Artificial Analysis Intelligence Index (BenchLM.ai snapshot, data verified July 23). Arena Elo = Arena.ai (formerly LMArena) text leaderboard; models in the pool fewer than ~2 weeks have low vote counts and scores will keep shifting — marked with †; some figures are carried from the June/early-July snapshot where no fresher data was found (marked ‡). SWE-bench Pro column throughout (not comparable to Verified scores — see footnotes for models where only Verified is published). "—" = not publicly confirmed or not yet evaluated.

Model Org Open? Arena Elo (Text) AA Index GPQA Dia SWE-bench Pro Terminal-Bench 2.1 Context $ / M in/out
Claude Fable 5 Anthropic No ~1,525‡ 59.9 (#1) 80.3% 86.0% 1M $10 / $50
GPT-5.6 Sol OpenAI No ~1,465 †‡ 58.9 (#2) 94.1% 64.6% 88.8% / 91.9%§ 1M+ $5 / $30
Kimi K3 Moonshot AI Yes* 1,486 † 57.1 (#3) 93.5% — (67.5% DeepSWE) 88.3% 1M $3 / $15
Claude Opus 4.8 Anthropic No ~1,510‡ 55.7 (#4) 69.2% ~85.0% 1M $5 / $25
GPT-5.6 Terra OpenAI No 55.0 (#5) 63.4% 87.4% 1M+ $2.50 / $15
Grok 4.5 xAI No ~1,496‡ 53.8 (#7) 93.0% 64.7% 500K $2 / $6
Gemini 3.6 Flash Google No 50.0 90.4% 58.7% (78% Verified) 1M $1.50 / $7.50
GLM-5.2 Zhipu / Z.ai Yes (MIT) 51.1 91.2% 62.1% 82.7% 1M $1.40 / $4.40
DeepSeek V4-Pro DeepSeek Yes (MIT) ~1,410‡ 44.0 — (80.6% Verified) 1M $0.44 / $0.87
Claude Sonnet 5 Anthropic No 53.4 (#9) 63.2% 80.4% 1M $2 / $10¶

*Kimi K3: API available; weights still pending as of July 24 — expected July 27 under Modified MIT. †Arena Elo: models that entered the pool within the last ~2–3 weeks have too few votes for stable rankings. ‡Arena Elo: no update found this week beyond the July 17 edition's figures; treat as directional, not current-day exact. §GPT-5.6 Sol (standard): 88.8%; Sol Ultra (4-agent Codex mode): 91.9%. ¶Claude Sonnet 5 at introductory pricing through August 31, 2026; then $3 / $15.

Reading the table: The top of the board is frozen — Fable 5, Sol, and K3 hold the same AA Index #1–#3 they held last week, and no new model challenged that order. The one real newcomer, Gemini 3.6 Flash, lands respectably (AA Index 50, above the ~31 category average for its tier) but is a mid-tier Flash model, not a Pro-class challenger — Google still has no answer at the frontier while it waits on Gemini 3.5 Pro. On pure GPQA Diamond, Grok 4.5's 93.0% edges past Kimi K3 (93.5% — statistically close) into the crowded 90%+ tier alongside Sol (94.1%) and GLM-5.2 (91.2%); the benchmark remains near-saturated at the top. DeepSeek V4-Pro's AA Index (44.0, freshly sourced this week) is notably lower than GLM-5.2's (51.1) despite similar GPQA figures reported elsewhere — a reminder that the AA Index is a composite across agentic, coding, and scientific tasks, not just knowledge recall, and DeepSeek's agentic/coding sub-scores are weaker. Fable 5's SWE-bench Pro lead (80.3%, ~11 points over Opus 4.8) remains the single clearest argument for its $10/$50 premium.


Benchmark & Leaderboard Movement

  • White House accuses Moonshot AI of distilling Anthropic's Fable to build Kimi K3 (July 22): OSTP director Michael Kratsios said the administration has evidence Moonshot ran a covert, industrial-scale distillation platform against US models — including Anthropic's Fable — while rotating access methods to avoid detection, and separately accessed banned Nvidia GB300 chips via Thailand. Treasury is now threatening sanctions. This is the first time a senior US official has named a specific Chinese lab and a specific American model in a distillation accusation, and it directly implicates the model (Kimi K3) that displaced Claude Fable 5 atop Arena's Frontend Code leaderboard just six days earlier. Kimi K3's benchmark standing hasn't moved, but its provenance is now a live diplomatic dispute — worth watching ahead of the July 27 open-weights release.
  • AA Intelligence Index top 3 holds steady: Claude Fable 5 (59.9), GPT-5.6 Sol (58.9), and Kimi K3 (57.1) are unchanged from the July 17 snapshot to the July 23 data pull — the first quiet week at the top since the index's post-June-launch churn.
  • Gemini 3.6 Flash enters mid-pack: At AA Index 50.0, Gemini 3.6 Flash slots just behind GLM-5.2 (51.1) and ahead of GPT-5.6 Luna (51.2 — essentially tied) — a Flash-tier model now benchmarking in open-weight-frontier territory, though still well off the top-3 cluster.
  • Qwen3.8-Max Preview: claim without evidence. Alibaba's "second only to Fable 5" claim has zero published benchmark backing as of July 23 — no technical report, no Artificial Analysis listing, no third-party eval. This is the loudest unverified claim of the week; treat it as marketing until real numbers land.
  • DeepSeek V4 migration deadline hits today: The deepseek-chat/deepseek-reasoner legacy endpoint names — flagged as retiring in last week's edition — go dead today, July 24, at 15:59 UTC, with no grace period. Both names have been silently routed to deepseek-v4-flash during the grace window; anything still calling the old names needs to be on deepseek-v4-pro or deepseek-v4-flash before the cutoff.
  • Grok 4.5 third-party benchmarks firm up: Independent coverage this week puts Grok 4.5 at 93.0% GPQA and 64.7% SWE-bench Pro — both now competitive with GPT-5.6 Sol's figures (94.1% / 64.6%), closing the gap from the sparse, mostly-self-reported picture in the July 8 edition.

Analysis

For agentic coding, nothing dislodged Claude Fable 5 this week — its 80.3% SWE-bench Pro score is still the clearest reason to pay $10/$50. Grok 4.5's newly-firmed 64.7% now edges out GPT-5.6 Sol (64.6%) in that specific metric, which is worth noting for teams already on xAI's stack, though the two are functionally tied. For reasoning/knowledge work, the 90–94% GPQA Diamond tier is crowded and close (Sol 94.1%, Grok 4.5 93.0%, Kimi K3 93.5%, GLM-5.2 91.2%, Gemini 3.6 Flash 90.4%) — pick on price and context rather than raw score. For cheap-and-fast production, Gemini 3.6 Flash ($1.50/$7.50) is a genuinely new, well-priced mid-tier option that beats its predecessor on cost while landing a competitive AA Index score; DeepSeek V4-Pro ($0.44/$0.87, MIT) remains the extreme-budget, self-hostable floor.

The open-vs-closed gap is unchanged from last week: Kimi K3 (open-pending) still sits between Sol and Opus 4.8 on the AA Index, and GLM-5.2 remains the most mature fully-open option in daily use. But this week is a reminder that "open" is getting geopolitically loaded — the White House's distillation accusation against Moonshot, arriving days before K3's promised weight release, means builders evaluating K3 for self-hosting should watch the July 27 date for both the weights themselves and any US regulatory response.


Sources

More from News