← September 2026
News 2026-09-25

AI Model & Benchmark Watch — September 25, 2026

Anthropic finally broke the three-way tie at the top of the leaderboard, OpenAI cut its mid-tier prices in half, Grok 4.7 arrived a week late with 40% more parameters and almost nothing to show for…

AI Model & Benchmark Watch — September 25, 2026

AI Model & Benchmark Watch — September 25, 2026

Anthropic finally broke the three-way tie at the top of the leaderboard, OpenAI cut its mid-tier prices in half, Grok 4.7 arrived a week late with 40% more parameters and almost nothing to show for it, and an open-weight model from Xiaomi's phone division out-scored it anyway.

Overview

This was the busiest week the frontier has had in a month: five models shipped in five days. Anthropic's Claude Opus 5.5 (September 22) jumped the Artificial Analysis Intelligence Index to 58, five points clear of the 53-53 tie that had held the top spot since Fable 5.1 and GPT-6 Astra landed. OpenAI answered the same day with GPT-6 Sol and GPT-6 Luna, its everyday and high-volume tiers, at roughly half of what the previous generation charged. Grok 4.7 finally shipped September 21, a week past Musk's own revised date, packing 2.1 trillion parameters against Grok 4.6's 1.5 trillion — and gained two points on the Index for the trouble. The open-weight story stole some of that thunder: Xiaomi's MiMo-V2.6-Pro, released the same day as Opus 5.5, took the open-weight Index lead at 46, edging past GLM-5.3 and Kimi K3 and landing in a flat tie with Grok 4.7, a closed model xAI just spent a week hyping.

New & Updated Models (this week)

Commercial / closed

Claude Opus 5.5 — Anthropic (September 22, 2026)

Who / License: Anthropic; closed, API and product access only.

What's notable: Opus 5.5 is Anthropic's first model to break away from the 53-53 tie that had sat atop the Artificial Analysis Index for weeks, scoring 58 — the first time in months a new model has opened real daylight over second place. It leads Terminal-Bench 4.0 at 59.6%, tied there with GPT-6 Astra at xhigh effort, and Anthropic's own reporting puts SWE-bench Pro at 89.9%, a vendor number that hasn't been independently reproduced yet. Context stays at 1M tokens, and pricing actually dropped: $4/$20 per million tokens, down from Opus 5's $5/$25, with cached reads at $0.20. Anthropic is also promising the model writes less "Claudish" prose, per its own release notes — a claim worth revisiting once people have used it for a few weeks.

Source: Vellum — Claude Opus 5.5 Benchmarks Explained · VentureBeat — Anthropic releases Claude Opus 5.5 · Artificial Analysis — Claude Opus 5.5 · The Decoder — Claude Opus 5.5 matches Fable 5.1 at 40% lower cost

GPT-6 Sol and GPT-6 Luna — OpenAI (September 22, 2026)

Who / License: OpenAI; closed, ChatGPT and API.

What's notable: These are OpenAI's mid-tier and budget-tier refreshes, built with the same training recipe as GPT-6 Astra. Sol drops to $2/$10 per million tokens (from GPT-5.6 Sol's $4/$20), scores 48 on the Intelligence Index, 64.6% on SWE-bench Pro, and carries a 1.05M-token context window. Luna undercuts everything in the field at $0.10/$0.50 per million tokens, scoring 37 on the Index — OpenAI says it makes "about half as many mistakes" as its predecessor, a claim from OpenAI's own release post that hasn't been independently checked.

Source: TechCrunch — OpenAI launches GPT-6 Sol and Luna · The New Stack — OpenAI cuts token prices in half · Artificial Analysis — GPT-6 Sol and Luna push the cost efficiency frontier · OpenAI — GPT-6 Sol model card

Grok 4.7 — xAI (September 21, 2026)

Who / License: xAI; closed, available via Grok Build, the Grok API, Cursor, and third-party routers.

What's notable: Grok 4.7 finally landed after xAI missed its own September 12 target, and Musk had already pre-graded it against Opus 5.0 on September 14 before anyone outside the company had run it. The headline spec is 2.1 trillion parameters, a 40% jump over Grok 4.6's 1.5 trillion, at the same $2/$6 per-million-token price (rising to $4/$12 above 200K tokens) and the same 500K context window. The gain from all that extra weight: two points on the Intelligence Index, from 44 to 46. On Terminal-Bench 4.0 it scores 26%, badly behind GPT-6 Astra's 60% and Claude Fable 5.1's 55% — and behind DeepSeek V4.1 Flash's 27%, a model a fraction of its price. The Decoder's framing was blunt: Grok 4.7's pricing looks like a Chinese lab's, and so does its benchmark position relative to the Western frontier.

Source: MarkTechPost — SpaceXAI Releases Grok 4.7 · The Decoder — xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap · Artificial Analysis — Benchmarking Grok 4.7 · llm-stats — Grok 4.7 Release: Same Price, Longer Horizons

Open-weight

MiMo-V2.6-Pro and MiMo-V2.6-Flash — Xiaomi (September 22, 2026)

Who / License: Xiaomi; open weights on Hugging Face under MIT.

What's notable: Xiaomi's phone-and-hardware arm shipped the new open-weight Index leader on the same day Anthropic and OpenAI both released flagships, and it didn't get buried: MiMo-V2.6-Pro scores 46 on the Artificial Analysis Index, ahead of GLM-5.3 and Kimi K3's 44 and tied with Grok 4.7. Pro runs 1.02 trillion parameters with 42 billion active per token; Flash runs 310 billion total with 15 billion active. Both are natively multimodal (text, image, audio, video in; text out, up to 128K tokens), share a 1M-token context window, and undercut the field on price — $0.435/$0.87 per million tokens for Pro, $0.14/$0.28 for Flash. Xiaomi also released a 9B distilled model, a technical report, and more than 7,000 RL environments used in training. The benchmark table Xiaomi published (DeepSWE v1.1, AutomationBench, Terminal-Bench 2.1, CyberGym) is vendor-run and uses an older Terminal-Bench version, so it isn't directly comparable to the 4.0 numbers elsewhere in this issue — treat it as a claim, not a verified result.

Source: VentureBeat — Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model · TechNode — Xiaomi open-sources MiMo-V2.6 models · SiliconANGLE — Xiaomi introduces MiMo-V2.6 series

Also this week: Alibaba previewed Qwen 4 at its Apsara Conference on September 22 — four tiers (Max, Flash, Plus, 27B), all still in training, with no release date, price, context window, or public benchmark yet. Worth watching, not yet worth a table entry.

Source: Yotta Labs — Qwen 4: Release Date, What's Confirmed · Pandaily — Alibaba Puts Qwen4 Family Into Training


Head-to-Head — Current Frontier

State of play as of September 25, 2026. AA Index is Artificial Analysis's v4.3.2, a minor calibration update from September 19 (re-anchoring GDPval-AA's Elo scale, refining fitting methodology) — not the kind of rescoring that invalidates comparisons to last week, unlike the Terminal-Bench overhaul on September 7. Terminal-Bench 4.0 figures below come from Artificial Analysis's own leaderboard and a September 22 Decoder comparison of the same three flagships; effort settings vary by model (noted where known), so treat close scores as roughly tied rather than precisely ranked. SWE-bench Verified stays frozen at the values archived May 4, 2026 as saturated; the Pro variant is still moving and figures marked * are vendor-reported, not independently reproduced. LMArena Elo is Arena.ai's text leaderboard — new releases take weeks to accumulate enough votes to rank, so several of this week's launches show "—" there.

Model Org Open? AA Index (v4.3.2) LMArena Elo GPQA Diamond SWE-bench (Verified/Pro) Terminal-Bench 4.0 Context $ / Mtok in/out
Claude Opus 5.5 Anthropic No 58 — (too new to rank) — (not reported) 89.9% (Pro)* 59.6% 1M $4 / $20
Claude Fable 5.1 Anthropic No 53 1498 92.6% 81.2% (Pro) 55% 1M $10 / $50
GPT-6 Astra OpenAI No 53 — (Code Arena: WebDev #1, 1800) 96.1% — 60% 1M $10 / $50
Claude Opus 5 Anthropic No 51 1493 84.1% 97.00% (Verified, frozen) / 79.2% (Pro) 51.8% 1M $5 / $25
Muse Spark 1.3 Meta No (weights undecided) 48 1493 — — — 1M $1.25 / $4.25
GPT-6 Sol OpenAI No 48 — — 64.6% (Pro) 43%* 1.05M $2 / $10
Grok 4.7 xAI No 46 — — — 26% 500K $2 / $6 (<200K)
MiMo-V2.6-Pro Xiaomi Yes (MIT) 46 — (too new to rank) — — — (2.1 self-reported)* 1M $0.44 / $0.87
GLM-5.3 Zhipu / Z.ai Yes (custom license, MaaS clause) 44 — — — — 1M $1.40 / $4.40
Kimi K3 Moonshot AI Yes (custom license) 44 — 93.5% 93.40% (Verified) — 1.05M $3 / $15
Gemini 3.8 Flash Google DeepMind No 41 1493 95.3% — — 1M $0.75 / $3.75
DeepSeek V4.1 Flash DeepSeek Yes (MIT) 39 — 90.9%* — 27% 1M $0.15 / $0.60

Reading the table: For the first time since Fable 5.1 and Astra tied at 53, there's real separation at the top — Opus 5.5's 58 is a five-point lead that no other model has approached this week. Astra still holds GPQA Diamond outright at 96.1%, and its Code Arena: WebDev win nudged up to 1,800 points, ahead of Fable 5.1 Max's 1,758. The open-weight ceiling moved for the first time in weeks too: MiMo-V2.6-Pro's 46 is two points above GLM-5.3 and Kimi K3, and it landed in a flat tie with Grok 4.7 — an open model from a phone maker matching a hyped closed release, on Index score if nothing else. On price, DeepSeek V4.1 Flash is still cheapest per token at $0.15/$0.60, but MiMo-V2.6-Pro undercuts every closed model in the table by a wide margin while scoring competitively, which is the more interesting price/performance story this week.


Benchmark & Leaderboard Movement

  • Opus 5.5 broke the three-way logjam at the top of the AA Index. Fable 5.1 and GPT-6 Astra had been tied at 53 for weeks; Opus 5.5 opened a five-point gap at 58, the first clear #1 since that tie formed.
  • MiMo-V2.6-Pro is the new open-weight Index leader at 46, up from GLM-5.3 and Kimi K3's 44 — a real, if modest, move in the open ceiling, and it happened the same day two closed flagships launched.
  • Grok 4.7's parameter count grew 40% for a two-point Index gain. The model went from 1.5T to 2.1T parameters and from an Index score of 44 to 46, while falling well behind on Terminal-Bench 4.0 (26% vs. Astra's 60% and Fable 5.1's 55%) — one of the largest hype-to-benchmark gaps this beat has tracked.
  • DeepSeek V4.1 Flash quietly beat Grok 4.7 on Terminal-Bench 4.0 (27% vs. 26%) despite costing a fraction as much, a detail buried in the same Decoder piece that broke down Grok 4.7's numbers.
  • The AA Index moved to v4.3.2 on September 19, a calibration-only update (re-anchoring GDPval-AA's Elo scale to DeepSeek V4.1 Flash at 1600, refined fitting for GDPval-AA and AA-Briefcase) rather than a benchmark swap — scores are comparable to last week's v4.3, unlike the Terminal-Bench 2.1-to-4.0 jump on September 7.
  • Qwen 4 is officially in training, previewed in four tiers at Alibaba's Apsara Conference on September 22, with Qwen 4.5 and Qwen 5 pitched on the roadmap at 5–10 trillion parameters — none of it shipping yet.

Analysis

For agentic coding, Opus 5.5 now leads on Terminal-Bench 4.0 at 59.6% (tied with GPT-6 Astra) and claims 89.9% on SWE-bench Pro by its own harness — worth a grain of salt until someone outside Anthropic reruns it. For reasoning, GPT-6 Astra's 96.1% on GPQA Diamond remains unmatched. For cheap-and-fast, MiMo-V2.6-Pro just became the most interesting option in the field: open weights, a 1M context window, native multimodality, and pricing under a dollar per million tokens combined, all while beating the two established open-weight leaders. For open-vs-closed, the gap actually widened this week even as the open ceiling rose — Opus 5.5's jump to 58 outpaced MiMo's two-point gain to 46, so the frontier's lead grew from 9 points to 12 in the same seven days. And for anyone tempted by Grok 4.7's parameter count: bigger clearly isn't better here, not against a field where a Chinese phone company's open model just matched it on the composite score at a fraction of the price.


Sources

More from News