← July 2026
News 2026-07-17

AI Model & Benchmark Watch — July 17, 2026

Moonshot AI's Kimi K3 — 2.8 trillion parameters, the world's largest open-weight model ever — debuted on July 16 at #3 on the Artificial Analysis Intelligence Index and seized the #1 spot on Arena's…

AI Model & Benchmark Watch — July 17, 2026

AI Model & Benchmark Watch — July 17, 2026

Moonshot AI's Kimi K3 — 2.8 trillion parameters, the world's largest open-weight model ever — debuted on July 16 at #3 on the Artificial Analysis Intelligence Index and seized the #1 spot on Arena's Frontend Code Arena, while Google fumbled Gemini 3.5 Pro's third consecutive deadline miss.

Overview

This week's story is Kimi K3: Moonshot AI dropped a 2.8T-parameter open-weight model on July 16 that immediately entered the top 3 of the Artificial Analysis Intelligence Index (57.1, behind only Fable 5 at 59.9 and GPT-5.6 Sol at 58.9) and knocked Claude Fable 5 off the #1 spot on Arena's Frontend Code leaderboard with a 1,679 Elo score. Full downloadable weights are promised by July 27, making it the first "open 3T-class" model. At the other end of the scale, PrismML shipped Bonsai 27B on July 14 — a 1-bit quantization of Qwen3.6-27B compressed to 3.9 GB that runs at 11 tokens/second on an iPhone 17 Pro under an Apache 2.0 license.

The week's notable non-event: Google DeepMind missed its July 17 target for Gemini 3.5 Pro — now the third consecutive deadline slip since the June I/O commitment — as Bloomberg reported the rebuilt model still falls short of internal coding benchmarks. Google is now reportedly considering a stopgap "Gemini 3.6 Flash" release while Pro continues tuning. Prediction markets have moved to August 7 at 73%. Meanwhile, DeepSeek graduated its V4 family from preview (since April 24) to official stable release this week, with old API endpoints (deepseek-chat, deepseek-reasoner) retiring July 24.

New & Updated Models (July 10–17)

Kimi K3 — Moonshot AI (July 16, 2026)

Who / License: Moonshot AI (China); open-weight pending — API live now, full weights releasing by July 27 under an expected Modified MIT license (same family as K2 series).

What's notable: Kimi K3 is a 2.8-trillion-parameter native multimodal Mixture-of-Experts model built on Kimi Delta Attention (KDA, a hybrid linear-attention mechanism) and Attention Residuals, activating 16 of 896 experts per token. Moonshot bills it as "the world's first open 3T-class model." Context window: 1M tokens. Reasoning: single effort level (max) at launch, with lower-effort tiers planned post-release.

Benchmark highlights (self-reported by Moonshot):

  • GPQA Diamond: 93.5% — the strongest published open-weight result on that benchmark at launch
  • Terminal-Bench 2.1: 88.3% — trails only GPT-5.6 Sol (88.8%) and Sol Ultra (91.9%)
  • BrowseComp: 91.2% — best published score on this benchmark at release
  • Humanity's Last Exam (with tools): 56.0%
  • MCP Atlas: 84.2%
  • DeepSWE: 67.5%

Third-party leaderboard standings (July 16–17):

  • Artificial Analysis Intelligence Index: 57.1 (#3, up from outside top 10 as K2.6)
  • Arena.ai Frontend Code Arena: #1 at 1,679 Elo — surpassing Claude Fable 5 (1,631) and GPT-5.6 Sol (1,618), ranking first in 6 of 7 Frontend domains
  • Arena.ai Text: #9 at 1,486 Elo — up from K2.6's #38 position

Pricing: $3.00 / M cache-miss input · $0.30 / M cached input · $15.00 / M output. Context: most expensive Chinese lab model released to date, but significantly cheaper per task than Opus 4.8 (~$0.94 per task vs. $1.80 per task).

Caveat: Until downloadable weights ship July 27, Artificial Analysis classifies K3 as proprietary (no self-hosting verification possible). Check the official model card before building commercial products on the expected Modified MIT.

Source: VentureBeat — Kimi K3 largest open-source model ever · Simon Willison — Kimi K3 and the pelican benchmark · Kimi K3 Tech Blog · OpenRouter — Kimi K3 pricing · Arena.ai on X — K3 #1 Frontend · CryptoBriefing — Kimi K3 dethroning Claude and GPT


PrismML Bonsai 27B — PrismML (July 14, 2026)

Who / License: PrismML (Caltech startup); Apache 2.0 — free to download and deploy.

What's notable: Bonsai 27B is a 1-bit and 1.58-bit ternary compression of Qwen3.6-27B. The 1-bit binary variant reaches 3.9 GB memory footprint — small enough to run in iPhone 17 Pro's RAM at 11 tokens/second while retaining ~90% of full-precision performance across 15 benchmarks. This is not a new pretrained model but a quantization artifact; performance on individual tasks will vary from the base. Apple is reportedly evaluating the compression technology for future on-device AI features.

Source: 9to5Mac — PrismML Bonsai 27B iPhone · MarkTechPost — Bonsai 27B · AlphaSignal — Bonsai 27B 3.9 GB · PrismML announcement


Gemini 3.5 Pro — Google DeepMind (Not Released — Third Deadline Miss)

Who / License: Google DeepMind; closed.

What happened: Gemini 3.5 Pro missed its July 17 general-availability target — the third consecutive slip since the original June I/O commitment (June GA → June 30 → July 17 → unknown). Bloomberg (July 16) reported the rebuilt model still falls short of internal coding benchmarks, with Google "taking time to improve capabilities, particularly in coding." Earlier reporting described hallucinations and inconsistent outputs as unresolved issues. Google has registered model names including Gemini 3.6 Flash and Gemini 3.5 Flash Light, widely read as stopgap releases to bridge the competitive gap while Pro continues its third rebuild cycle. As of this writing, no model card, pricing page, or API listing is available for Gemini 3.5 Pro.

Rumored specs (unconfirmed, from third-party sources only): 2M-token context window, "Deep Think" reasoning mode, pricing near $1.25/$10 per million tokens at standard tiers and $15/$60 for Deep Think; Deep Think gated to the $250/month Ultra subscription.

What to do: Do not plan production workloads around a July date. Prediction markets now put August 7 at 73%. Watch for an official Google model card; every number before that is a leak.

Source: Bloomberg — Gemini launch delayed, falls short of internal goals · TechTimes — Gemini 3.5 Pro misses third deadline · Geeky Gadgets — Stopgap Gemini 3.6 Flash · 9to5Google — Gemini 3.5 Pro delays, coding issues


DeepSeek V4 — Official Stable GA (July 17, 2026)

Who / License: DeepSeek; MIT — full open weights.

What happened: DeepSeek V4 Pro and V4 Flash graduated from preview (available since April 24) to official stable release on July 17 with stated performance and feature improvements. This is not a new model but the graduation milestone. Simultaneously, DeepSeek announced that legacy endpoints (deepseek-chat and deepseek-reasoner) will be permanently retired on July 24 at 15:59 UTC — any production code calling those endpoints must migrate to deepseek-v4-pro or deepseek-v4-flash before that date.

Specs (unchanged from preview): V4 Pro: 1.6T-parameter MoE, ~49B active parameters per token, $0.44/$0.87 per M tokens, 1M context. V4 Flash: 284B total, ~13B active, fast/cheap workhorse.

Source: DeepSeek API Docs — V4 Preview Release · ExplainX — DeepSeek V4 official release pricing · TechTimes — Gemini 3.5 Pro vs DeepSeek July 24 deadline


Mistral Robostral Navigate — Mistral AI (July 8, 2026) carried from prior week

Who / License: Mistral AI; open weight (Apache 2.0).

What's notable: An 8B embodied-navigation model trained on ~400K simulation trajectories across 6,000 scenes. Takes a single RGB camera feed plus plain-language instructions (no LiDAR, no depth sensors) and navigates robots through unseen environments at 76.6% success rate on R2R-CE validation — 9.7 points ahead of the prior best single-camera system and 4.5 points ahead of systems using depth or multiple cameras. Hardware-agnostic: runs on wheeled, legged, and flying robots. This was released July 8, after the prior edition's close, and warrants a mention for builders in robotics/embodied AI.

Source: Bloomberg — Mistral releases robotics model · Pulse2 — Robostral Navigate launch · Technology.org — Robostral Navigate overview


Head-to-Head — Current Frontier

The table below reflects the state of play as of July 17, 2026. AA Index = Artificial Analysis Intelligence Index (scores from BenchLM.ai July 16 update, aligned with Artificial Analysis data). Arena Elo = Arena.ai (formerly LMArena); models in the arena fewer than ~2 weeks have low vote counts and scores will shift materially — marked with †. SWE-bench column = Pro variant throughout (not comparable to Verified scores). "—" = not publicly confirmed or not yet evaluated.

Model Org Open? Arena Elo (Text) AA Index GPQA Dia SWE-bench Pro Terminal-Bench 2.1 Context $ / M in/out
Claude Fable 5 Anthropic No 59.9 (#1) 80.3% 86.0% 1M $10 / $50
GPT-5.6 Sol OpenAI No ~1,465 † 58.9 (#2) 94.1% 64.6% 88.8% / 91.9%‡ 1M+ $5 / $30
Kimi K3 Moonshot AI Yes* 1,486 † 57.1 (#3) 93.5% 67.5 (DeepSWE) 88.3% 1M $3 / $15
Claude Opus 4.8 Anthropic No ~1,510 55.7 (#4) 69.2% ~85.0% 1M $5 / $25
GPT-5.6 Terra OpenAI No 55.0 (#5) 63.4% 87.4% 1M+ $2.50 / $15
Grok 4.5 xAI No 53.8 (#7) 500K $2 / $6
Gemini 3.1 Pro Google No ~1,505 ~46 94.3% 54.2% 1M $2 / $12
GLM-5.2 Zhipu / Z.ai Yes (MIT) 51 91.2% 62.1% 81.0% 1M ~$1.18 avg
DeepSeek V4-Pro DeepSeek Yes (MIT) ~1,410 90.1% 1M $0.44 / $0.87
Claude Sonnet 5 Anthropic No 53.4 (#9) 63.2% 80.4% 1M $2 / $10§

*Kimi K3: API available; weights releasing July 27 under expected Modified MIT — classified proprietary by AA until weights ship.
†Arena Elo: models entered the pool within the last 2 weeks have too few votes for stable rankings — scores will move significantly.
‡GPT-5.6 Sol (standard): 88.8%; Sol Ultra (4-agent Codex mode): 91.9%.
§Claude Sonnet 5 at introductory pricing through August 31, 2026; then $3 / $15.

Reading the table: The biggest structural change this week is Kimi K3 forcing its way into the #3 AA Index slot — the first time a Chinese lab has placed a model in the top 3 of that index. Kimi K3 is competitive with GPT-5.6 Sol on Terminal-Bench (88.3% vs. 88.8%) and GPQA Diamond (93.5% vs. 94.1%), and its DeepSWE score (67.5%) trails Sol's SWE-bench Pro (64.6%) by a different benchmark's metric — the two aren't directly comparable but both point to the same tier. The open-vs-closed gap at the frontier has closed materially: K3 at $3/$15 is cheaper than Sol ($5/$30) and approaching Fable 5 territory on agentic capability, all under an open license. Claude Fable 5 retains the definitive lead on SWE-bench Pro (80.3%, 15+ points clear of the field) — still the hardest coding benchmark advantage to challenge. On the value tier, DeepSeek V4-Pro now at official stable ($0.44/$0.87, MIT) remains the extreme-budget option for teams who can self-host.


Benchmark & Leaderboard Movement

  • AA Intelligence Index — Kimi K3 enters top 3: K3 debuted at 57.1 (#3 on BenchLM's July 16 update), pushing Claude Opus 4.8 (55.7) to #4 — the first time a Chinese lab model has entered the AA Intelligence Index top 3. The gap from #1 Fable 5 (59.9) to #3 K3 (57.1) is 2.8 points, the tightest top-3 spread since the index launched.
  • Arena Frontend Code Arena — new #1: Kimi K3 debuted at 1,679 Elo, displacing Claude Fable 5 (1,631) and GPT-5.6 Sol (1,618). K3 is #1 in 6 of 7 frontend domains; Fable 5 holds only Gaming. This is a dramatic 17-place jump from K2.6's previous #18 rank.
  • Arena Text — K3 lands at #9: At 1,486 Elo (3,026 votes), K3 has vaulted from K2.6's #38 text rank to #9. With low vote counts, this will shift; however, the directional signal is unusually strong given the sample size.
  • GPQA Diamond approaching saturation: Sakana Fugu-Ultra leads at 95.5%, while five models cluster in the 93–94.3% range (GPT-5.6 Sol 94.1%, Gemini 3.1 Pro 94.3%, Kimi K3 93.5%). Improvement on this benchmark is increasingly marginal; leaderboard watchers are moving focus to agentic benchmarks (BrowseComp, Terminal-Bench, SWE-bench Pro) as the meaningful differentiators.
  • SWE-bench Pro — Fable 5 moat deepens relative to the field: The July 10 edition noted Fable 5 at 80.3%; BenchLM coding data now shows 81.7% on the Lite variant, with GPT-5.6 Sol at 74.1% on the same variant. The gap is wider than raw Pro scores suggested.
  • Gemini 3.1 Pro slides in relevance: With Kimi K3 now above it on the AA Index and Gemini 3.5 Pro pushed to at least August, Google's position on the frontier leaderboard is eroding — Gemini 3.1 Pro's AA Index score (~46) puts it behind every major competitor except budget models.
  • DeepSeek V4 stable: V4 Pro's graduation from preview removes the "preview" caveat for production use. Teams using old deepseek-chat / deepseek-reasoner API endpoints have one week (until July 24) to migrate before forced cutoff.

Analysis

For agentic coding and software engineering, the picture sharpens again: Claude Fable 5 ($10/$50) is still the only model with an 80%+ SWE-bench Pro score and the clearest agentic benchmark advantage. But this week's Kimi K3 complicates the calculus — at $3/$15 with weights arriving July 27, it offers AA Index #3 capability with the self-hosting option that all closed models lack. For teams that can run it, K3 is the first open-weight model that honestly belongs in the same conversation as GPT-5.6 Sol at a fraction of the cost.

For terminal and agent-style work, Sol Ultra (91.9% Terminal-Bench, multi-agent Codex) still leads the benchmark that matters most here; standard Sol (88.8%) and Kimi K3 (88.3%) are within noise of each other. For reasoning and scientific knowledge (GPQA Diamond), the benchmark is effectively saturated — the 94% tier contains three models at wildly different price points ($2/$12 for Gemini 3.1 Pro, $3/$15 for K3, $5/$30 for Sol). For cheap-and-fast production, DeepSeek V4-Pro at stable ($0.44/$0.87) and Claude Sonnet 5 ($2/$10 introductory through August 31) remain the anchors. For on-device / edge, PrismML Bonsai 27B is the week's most interesting new option for mobile-first builders.

The open-vs-closed gap at the frontier is at its lowest point to date: Kimi K3 sits between GPT-5.6 Sol and Claude Opus 4.8 on the AA Intelligence Index, runs on open weights (next week), and benchmarks competitively on GPQA and Terminal-Bench. The remaining closed-model advantage concentrates at one specific benchmark: SWE-bench Pro, where Claude Fable 5's 80.3% gap over the field ($10/$50 pricing) is the clearest argument for paying premium rates.


Sources

More from News