AI Model & Benchmark Watch — July 8, 2026
xAI launched Grok 4.5 today as its first self-described "Opus-class" model, GPT-5.6 Sol broke the Terminal-Bench record and sits one government approval away from broad release, and open-weight…
AI Model & Benchmark Watch — July 8, 2026
xAI launched Grok 4.5 today as its first self-described "Opus-class" model, GPT-5.6 Sol broke the Terminal-Bench record and sits one government approval away from broad release, and open-weight GLM-5.2 is quietly beating GPT-5.5 on long-horizon coding for ~1/6th the cost.
Overview
The week of July 8 is arguably the densest product week since Claude Fable 5 launched in early June. xAI shipped Grok 4.5 this morning — built on a 1.5-trillion-parameter V9 foundation co-trained with the Cursor coding editor (acquired by SpaceX for $60B in June), priced at $2/$6 per million tokens and immediately the default in Grok Build and Cursor. OpenAI's GPT-5.6 family (Sol, Terra, Luna) previewed June 26 but remains in a government-gated partner preview; prediction markets placed broad availability on July 9, and Sol's Terminal-Bench 2.1 score of 88.8% (91.9% in Ultra mode) is the highest published figure on that benchmark. On the open-weight side, Zhipu AI / Z.ai's GLM-5.2 — a 744B MIT-licensed MoE model released June 13 — continues to generate discussion after scoring 62.1% on SWE-bench Pro, beating GPT-5.5's 58.6% at a fraction of the API cost. Meanwhile, Claude Sonnet 5 (June 30) is now the default model for all Free and Pro claude.ai users, landing within a few points of Opus 4.8 on agentic coding benchmarks at introductory pricing of $2/$10.
New & Updated Models (this week / past ~7 days)
Grok 4.5 — xAI (July 8, 2026)
Who / License: xAI (SpaceXAI); closed, proprietary.
What's notable: Built on xAI's 1.5T-parameter V9 foundation and co-trained alongside the Cursor IDE (which SpaceX acquired for $60B). Released to Grok Build, Cursor (all plans), and the SpaceXAI console. Musk describes it as "Opus-class, but faster and more token-efficient." On the four benchmarks xAI published, Grok 4.5 beats Claude Opus 4.8 on DeepSWE 1.0 and Terminal-Bench 2.1, but trails it on DeepSWE 1.1 (by 6 points) and SWE-bench Pro (by 4.5 points). Notably scores #1 on Harvey's Legal Agent Benchmark. BenchLM notes it currently has only 6 published scores across 254 tracked benchmarks, so rankings are provisional. Context window: 500K tokens. Speed: ~80 tokens/second. Price: $2/$6 per million tokens.
Source: SpaceXAI Introducing Grok 4.5 · TechCrunch · Axios · BenchLM Grok 4.5 · LM Market Cap
GPT-5.6 Sol / Terra / Luna — OpenAI (previewed June 26, 2026; broad release pending)
Who / License: OpenAI; closed, proprietary.
What's notable: Three-tier family: Sol (flagship), Terra (balanced, GPT-5.5-quality at 2× lower cost), Luna (fastest/cheapest). At the request of the U.S. Department of Commerce, OpenAI staged rollout to ~20 trusted partner organizations first; broad availability was not live as of July 7 but is described as "in the coming weeks." Sol sets a new state-of-the-art on Terminal-Bench 2.1 at 88.8% (91.9% in Ultra/Codex multi-agent mode), ahead of Claude Fable 5's 83.4%. Context window: 1.5M tokens (up from GPT-5.5's 1M). Pricing (preview): Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million tokens.
Source: OpenAI GPT-5.6 Sol preview · ExplainX GPT-5.6 guide · TechTimes — METR risk flag
Claude Sonnet 5 — Anthropic (June 30, 2026)
Who / License: Anthropic; closed, proprietary.
What's notable: Anthropic's "most agentic Sonnet yet" — model ID claude-sonnet-5. Now the default model for Free and Pro users on claude.ai, and live in Claude Code, Cursor, VS Code, and GitHub Copilot. SWE-bench Pro: 63.2% (vs Opus 4.8's 69.2%, Sonnet 4.6's 58.1%); SWE-bench Verified: 85.2%; Terminal-Bench 2.1: 80.4%; BrowseComp single-agent: 84.7%; USAMO 2026: 79.5%. Supports adaptive thinking with selectable effort up to xhigh. Context: 1M tokens. Introductory pricing: $2/$10 per million tokens through August 31, 2026 (then $3/$15).
Source: DataCamp Sonnet 5 · CosmicJS Sonnet 5 · LLM Stats Sonnet 5 · Vellum Sonnet 5 benchmarks
Seed 2.1 Pro & Turbo — ByteDance (June 24, 2026)
Who / License: ByteDance; closed, proprietary.
What's notable: Two variants expanding the Seed 2.1 series — Pro (performance-focused) and Turbo (speed-optimized). Details are limited in public trackers; LLM Stats lists both as new tracked models but has not published full benchmark coverage yet.
Source: LLM Stats model updates
Open-Weight Highlights (recent weeks)
GLM-5.2 — Zhipu AI / Z.ai (June 13, 2026)
Who / License: Zhipu AI (Z.ai); open-weight, MIT license — fully permissive, commercially usable, self-hostable.
What's notable: 744B-parameter MoE (activates ~40B per token) with a usable 1M-token context window and an architectural innovation called IndexShare that reduces per-token FLOPs ~2.9× at full context. On SWE-bench Pro: 62.1% — ahead of GPT-5.5's 58.6% and Gemini 3.1 Pro's 54.2%. On MCP-Atlas (multi-tool agent workflows): 77.0, statistically level with Claude Opus 4.8's 77.8. Terminal-Bench 2.1: 81.0, within a few points of Opus 4.8's 85.0. GPQA Diamond: 91.2%. Weights on Hugging Face under MIT license. API: ~$1.18/M average (per LLM Stats). VentureBeat headlined it as beating GPT-5.5 at ~1/6th the cost.
Source: VentureBeat GLM-5.2 · Technology.org GLM-5.2 · The AI Rankings GLM-5.2
Kimi K2.7 Code — Moonshot AI (June 12, 2026)
Who / License: Moonshot AI; open-weight, Modified MIT license.
What's notable: 1T-parameter MoE (32B active), 256K context, multimodal (text + image via MoonViT). Reports +21.8% on Kimi Code Bench v2 over K2.6 and 30% reduction in thinking-token usage. Caveat: all published benchmarks are Moonshot proprietary suites — no independent SWE-bench/LiveCodeBench/GPQA scores published at release. API: $0.74/$3.50 per million tokens.
Source: MarkTechPost Kimi K2.7 · DevOps.com
Head-to-Head — Current Frontier
The table below covers the current top frontier models. Arena Elo figures are from a June 2026 snapshot (source: LocalAI Master) — models released after mid-June (Fable 5, Sonnet 5, Grok 4.5) haven't accumulated enough votes to rank reliably and are shown as "—". AA Index = Artificial Analysis Intelligence Index v3 (out of 170 models, higher = more intelligent). SWE-bench column uses the Pro variant where available (harder than Verified; scores not directly comparable across variants). Pricing is per 1M tokens, input/output; "—" means not publicly confirmed.
| Model | Org | Open? | Arena Elo | AA Index | GPQA Dia | SWE-bench Pro | Context | $/M in/out |
|---|---|---|---|---|---|---|---|---|
| Claude Mythos Preview | Anthropic | No (invite-only) | — | — | 94.6% | 77.8% | 1M | $25/$125 |
| Claude Fable 5 | Anthropic | No | — | 60 (#1/170) | — | 80.3% | 1M | $10/$50 |
| GPT-5.6 Sol | OpenAI | No (preview) | — | — | — | — | 1.5M | $5/$30 |
| Gemini 3.1 Pro Preview | No | ~1505 | 46 (#13/170) | 94.3% | 54.2% | 1M | $2/$12 | |
| Claude Opus 4.8 | Anthropic | No | ~1510 | — | — | 69.2% | 1M | ~$7.22 avg |
| Grok 4.5 | xAI | No | — | — | — | — | 500K | $2/$6 |
| Claude Sonnet 5 | Anthropic | No | — | — | — | 63.2% | 1M | $2/$10* |
| Qwen 3.7 Max | Alibaba | No (API) | ~1488 | — | 92.3% | — | 1M | $1.25/$3.75 |
| GLM-5.2 | Zhipu/Z.ai | Yes (MIT) | — | — | 91.2% | 62.1% | 1M | ~$1.18 avg |
| DeepSeek V4-Pro | DeepSeek | Yes (open) | ~1410 | — | 90.1% | — | 1M | $0.44/$0.87 |
*Claude Sonnet 5 introductory pricing through August 31, 2026; then $3/$15. †Arena Elo figures are approximate (~) from a June 2026 article snapshot; the leaderboard updates continuously. ‡Claude Opus 4.8 average price from LLM Stats; in/out split not confirmed.
Reading the table: Claude Fable 5 leads on the Artificial Analysis Intelligence Index (#1/170) and SWE-bench Pro (80.3%), while Claude Mythos Preview holds the highest GPQA Diamond score (94.6%) of any model in the table — but access is invitation-only at $25/$125. GPT-5.6 Sol's Terminal-Bench 2.1 record (88.8%; not in this table — no SWE-bench Pro figure yet) makes it the one to watch for terminal/agentic coding once it goes GA. Among open-weight models, GLM-5.2 is the standout: 91.2% GPQA and 62.1% SWE-bench Pro under an MIT license, beating GPT-5.5 on the coding benchmark at ~1/6th the cost. DeepSeek V4-Pro at $0.87/M output is the cheapest path to frontier-tier reasoning in the table.
Benchmark & Leaderboard Movement
- New record — Terminal-Bench 2.1: GPT-5.6 Sol Ultra (91.9%, multi-agent Codex mode) broke the prior record held by Claude Fable 5 (83.4%). Sol in standard config is 88.8%. This benchmark measures autonomous terminal task completion — the most agent-relevant eval at the frontier right now.
- Arena Elo cluster tightens: As of June 2026, the top 6 models on LMArena span only ~55 Elo points (Claude Opus 4.8 ~1510 to Grok 4.3 ~1496), the tightest spread on record — rendering marginal Elo differences nearly meaningless for practical use.
- Grok 4.5 debuts today (#1 on Harvey Legal): xAI's Harvey Legal Agent Benchmark claim is notable. Grok 4.5 also beats Opus 4.8 on DeepSWE 1.0 and Terminal-Bench 2.1 on xAI's internal runs. Independent third-party eval coverage is still sparse.
- Open-weight closes coding gap: GLM-5.2 (open, MIT) scores 62.1% SWE-bench Pro, sitting within 1.1 points of Claude Sonnet 5 (63.2%) — the best closed model at a far lower price point. This is arguably the biggest story in open-vs-closed parity this week.
- Anthropic's intelligent index lead: Claude Fable 5 holds AA Intelligence Index rank #1 of 170 models (score: 60); Gemini 3.1 Pro is #13 (score: 46). The gap between these two on the index is larger than any other adjacent pair in the top 15.
- Benchmark-integrity note — Kimi K2.7 Code: Moonshot AI's June 12 release shows dramatic gains on its own benchmarks (+21.8% Kimi Code Bench v2) but published no scores on standard suites (SWE-bench, LiveCodeBench, GPQA). Treat proprietary benchmark claims with caution until independent results emerge.
Analysis
For agentic coding, Claude Fable 5 is the ceiling on SWE-bench Pro (80.3%) and Artificial Analysis's composite index, but at $10/$50 it's expensive — Claude Sonnet 5 hits 63.2% at $2/$10 introductory and is live in Claude Code today, making it the best-value coding agent for most teams. Reasoning and knowledge (GPQA Diamond) is a three-way tie between Claude Mythos Preview (94.6%, invite-only), Gemini 3.1 Pro (94.3%, $2/$12), and Qwen 3.7 Max (92.3%, $1.25/$3.75) — Gemini and Qwen are the accessible picks. For cheap-and-fast, GPT-5.6 Luna ($1/$6, once GA) and Grok 4.1 Fast ($0.20/$0.50) are the frontrunners; DeepSeek V4-Flash ($0.14/$0.28, open-weight) undercuts them both. The open-vs-closed gap has effectively closed at the $1–2/M price tier for coding: GLM-5.2 and DeepSeek V4-Pro are genuine frontier-class open alternatives, not catch-up models.
Sources
- LLM Stats — AI model leaderboard and updates
- LLM Stats — AI Updates (July 2026)
- LLM Stats — Claude Sonnet 5 model page
- LocalAI Master — LMArena Leaderboard 2026 (Arena Elo snapshot)
- Arena AI official leaderboard
- Hugging Face — Arena Leaderboard (lmarena-ai)
- Artificial Analysis — Claude Fable 5 (AA Intelligence Index, pricing, speed)
- Artificial Analysis — Gemini 3.1 Pro Preview
- Anthropic — Claude Fable 5 benchmarks (morphllm.com summary)
- BenchLM — Claude Fable 5
- Vellum — Claude Fable 5 & Mythos 5 benchmarks
- CosmicJS — Claude Fable 5
- DataCamp — Claude Sonnet 5
- CosmicJS — Claude Sonnet 5
- Vellum — Claude Sonnet 5 benchmarks explained
- CodersEra — Claude Sonnet 5 launch guide
- Anthropic — Claude Mythos Preview cybersecurity assessment
- NxCode — Claude Mythos Preview
- LLM Stats — Claude Mythos Preview
- Claude Platform Docs — Models overview
- OpenAI — Previewing GPT-5.6 Sol
- ExplainX — GPT-5.6 Sol, Terra, Luna guide
- Coursiv — ChatGPT 5.6 release details
- SpaceXAI — Introducing Grok 4.5
- TechCrunch — Grok 4.5 launch
- Axios — Grok 4.5 scoop
- BenchLM — Grok 4.5
- LM Market Cap — Grok 4.5
- Roo (Beehiiv) — Grok 4.5 benchmark analysis
- Artificial Analysis — Grok 4.3
- OpenRouter — Gemini 3.1 Pro Preview
- LLM Stats — Gemini 3.1 Pro blog
- VentureBeat — GLM-5.2 beats GPT-5.5
- Technology.org — GLM-5.2 coding benchmarks
- The AI Rankings — GLM-5.2
- HuggingFace — DeepSeek-V4-Pro model card
- MorphLLM — DeepSeek V4 architecture & pricing
- DataCamp — DeepSeek V4
- OpenRouter — Qwen 3.7 Max
- CodersEra — Qwen 3.7 Max launch guide
- LLM Stats — Qwen 3.7 Max model page
- MarkTechPost — Kimi K2.7 Code release
- DevOps.com — Kimi K2.7 Code
- BenchLM — LLM Leaderboard July 2026
- Skycrumbs — AI models July 2026 tracker
More from News