AI Model & Benchmark Watch — September 25, 2026
Anthropic finally broke the three-way tie at the top of the leaderboard, OpenAI cut its mid-tier prices in half, Grok 4.7 arrived a week late with 40% more parameters and almost nothing to show for…
AI Model & Benchmark Watch — September 25, 2026
Anthropic finally broke the three-way tie at the top of the leaderboard, OpenAI cut its mid-tier prices in half, Grok 4.7 arrived a week late with 40% more parameters and almost nothing to show for it, and an open-weight model from Xiaomi's phone division out-scored it anyway.
Overview
This was the busiest week the frontier has had in a month: five models shipped in five days. Anthropic's Claude Opus 5.5 (September 22) jumped the Artificial Analysis Intelligence Index to 58, five points clear of the 53-53 tie that had held the top spot since Fable 5.1 and GPT-6 Astra landed. OpenAI answered the same day with GPT-6 Sol and GPT-6 Luna, its everyday and high-volume tiers, at roughly half of what the previous generation charged. Grok 4.7 finally shipped September 21, a week past Musk's own revised date, packing 2.1 trillion parameters against Grok 4.6's 1.5 trillion — and gained two points on the Index for the trouble. The open-weight story stole some of that thunder: Xiaomi's MiMo-V2.6-Pro, released the same day as Opus 5.5, took the open-weight Index lead at 46, edging past GLM-5.3 and Kimi K3 and landing in a flat tie with Grok 4.7, a closed model xAI just spent a week hyping.
New & Updated Models (this week)
Commercial / closed
Claude Opus 5.5 — Anthropic (September 22, 2026)
Who / License: Anthropic; closed, API and product access only.
What's notable: Opus 5.5 is Anthropic's first model to break away from the 53-53 tie that had sat atop the Artificial Analysis Index for weeks, scoring 58 — the first time in months a new model has opened real daylight over second place. It leads Terminal-Bench 4.0 at 59.6%, tied there with GPT-6 Astra at xhigh effort, and Anthropic's own reporting puts SWE-bench Pro at 89.9%, a vendor number that hasn't been independently reproduced yet. Context stays at 1M tokens, and pricing actually dropped: $4/$20 per million tokens, down from Opus 5's $5/$25, with cached reads at $0.20. Anthropic is also promising the model writes less "Claudish" prose, per its own release notes — a claim worth revisiting once people have used it for a few weeks.
Source: Vellum — Claude Opus 5.5 Benchmarks Explained · VentureBeat — Anthropic releases Claude Opus 5.5 · Artificial Analysis — Claude Opus 5.5 · The Decoder — Claude Opus 5.5 matches Fable 5.1 at 40% lower cost
GPT-6 Sol and GPT-6 Luna — OpenAI (September 22, 2026)
Who / License: OpenAI; closed, ChatGPT and API.
What's notable: These are OpenAI's mid-tier and budget-tier refreshes, built with the same training recipe as GPT-6 Astra. Sol drops to $2/$10 per million tokens (from GPT-5.6 Sol's $4/$20), scores 48 on the Intelligence Index, 64.6% on SWE-bench Pro, and carries a 1.05M-token context window. Luna undercuts everything in the field at $0.10/$0.50 per million tokens, scoring 37 on the Index — OpenAI says it makes "about half as many mistakes" as its predecessor, a claim from OpenAI's own release post that hasn't been independently checked.
Source: TechCrunch — OpenAI launches GPT-6 Sol and Luna · The New Stack — OpenAI cuts token prices in half · Artificial Analysis — GPT-6 Sol and Luna push the cost efficiency frontier · OpenAI — GPT-6 Sol model card
Grok 4.7 — xAI (September 21, 2026)
Who / License: xAI; closed, available via Grok Build, the Grok API, Cursor, and third-party routers.
What's notable: Grok 4.7 finally landed after xAI missed its own September 12 target, and Musk had already pre-graded it against Opus 5.0 on September 14 before anyone outside the company had run it. The headline spec is 2.1 trillion parameters, a 40% jump over Grok 4.6's 1.5 trillion, at the same $2/$6 per-million-token price (rising to $4/$12 above 200K tokens) and the same 500K context window. The gain from all that extra weight: two points on the Intelligence Index, from 44 to 46. On Terminal-Bench 4.0 it scores 26%, badly behind GPT-6 Astra's 60% and Claude Fable 5.1's 55% — and behind DeepSeek V4.1 Flash's 27%, a model a fraction of its price. The Decoder's framing was blunt: Grok 4.7's pricing looks like a Chinese lab's, and so does its benchmark position relative to the Western frontier.
Source: MarkTechPost — SpaceXAI Releases Grok 4.7 · The Decoder — xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap · Artificial Analysis — Benchmarking Grok 4.7 · llm-stats — Grok 4.7 Release: Same Price, Longer Horizons
Open-weight
MiMo-V2.6-Pro and MiMo-V2.6-Flash — Xiaomi (September 22, 2026)
Who / License: Xiaomi; open weights on Hugging Face under MIT.
What's notable: Xiaomi's phone-and-hardware arm shipped the new open-weight Index leader on the same day Anthropic and OpenAI both released flagships, and it didn't get buried: MiMo-V2.6-Pro scores 46 on the Artificial Analysis Index, ahead of GLM-5.3 and Kimi K3's 44 and tied with Grok 4.7. Pro runs 1.02 trillion parameters with 42 billion active per token; Flash runs 310 billion total with 15 billion active. Both are natively multimodal (text, image, audio, video in; text out, up to 128K tokens), share a 1M-token context window, and undercut the field on price — $0.435/$0.87 per million tokens for Pro, $0.14/$0.28 for Flash. Xiaomi also released a 9B distilled model, a technical report, and more than 7,000 RL environments used in training. The benchmark table Xiaomi published (DeepSWE v1.1, AutomationBench, Terminal-Bench 2.1, CyberGym) is vendor-run and uses an older Terminal-Bench version, so it isn't directly comparable to the 4.0 numbers elsewhere in this issue — treat it as a claim, not a verified result.
Source: VentureBeat — Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model · TechNode — Xiaomi open-sources MiMo-V2.6 models · SiliconANGLE — Xiaomi introduces MiMo-V2.6 series
Also this week: Alibaba previewed Qwen 4 at its Apsara Conference on September 22 — four tiers (Max, Flash, Plus, 27B), all still in training, with no release date, price, context window, or public benchmark yet. Worth watching, not yet worth a table entry.
Source: Yotta Labs — Qwen 4: Release Date, What's Confirmed · Pandaily — Alibaba Puts Qwen4 Family Into Training
Head-to-Head — Current Frontier
State of play as of September 25, 2026. AA Index is Artificial Analysis's v4.3.2, a minor calibration update from September 19 (re-anchoring GDPval-AA's Elo scale, refining fitting methodology) — not the kind of rescoring that invalidates comparisons to last week, unlike the Terminal-Bench overhaul on September 7. Terminal-Bench 4.0 figures below come from Artificial Analysis's own leaderboard and a September 22 Decoder comparison of the same three flagships; effort settings vary by model (noted where known), so treat close scores as roughly tied rather than precisely ranked. SWE-bench Verified stays frozen at the values archived May 4, 2026 as saturated; the Pro variant is still moving and figures marked * are vendor-reported, not independently reproduced. LMArena Elo is Arena.ai's text leaderboard — new releases take weeks to accumulate enough votes to rank, so several of this week's launches show "—" there.
| Model | Org | Open? | AA Index (v4.3.2) | LMArena Elo | GPQA Diamond | SWE-bench (Verified/Pro) | Terminal-Bench 4.0 | Context | $ / Mtok in/out |
|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | No | 58 | — (too new to rank) | — (not reported) | 89.9% (Pro)* | 59.6% | 1M | $4 / $20 |
| Claude Fable 5.1 | Anthropic | No | 53 | 1498 | 92.6% | 81.2% (Pro) | 55% | 1M | $10 / $50 |
| GPT-6 Astra | OpenAI | No | 53 | — (Code Arena: WebDev #1, 1800) | 96.1% | — | 60% | 1M | $10 / $50 |
| Claude Opus 5 | Anthropic | No | 51 | 1493 | 84.1% | 97.00% (Verified, frozen) / 79.2% (Pro) | 51.8% | 1M | $5 / $25 |
| Muse Spark 1.3 | Meta | No (weights undecided) | 48 | 1493 | — | — | — | 1M | $1.25 / $4.25 |
| GPT-6 Sol | OpenAI | No | 48 | — | — | 64.6% (Pro) | 43%* | 1.05M | $2 / $10 |
| Grok 4.7 | xAI | No | 46 | — | — | — | 26% | 500K | $2 / $6 (<200K) |
| MiMo-V2.6-Pro | Xiaomi | Yes (MIT) | 46 | — (too new to rank) | — | — | — (2.1 self-reported)* | 1M | $0.44 / $0.87 |
| GLM-5.3 | Zhipu / Z.ai | Yes (custom license, MaaS clause) | 44 | — | — | — | — | 1M | $1.40 / $4.40 |
| Kimi K3 | Moonshot AI | Yes (custom license) | 44 | — | 93.5% | 93.40% (Verified) | — | 1.05M | $3 / $15 |
| Gemini 3.8 Flash | Google DeepMind | No | 41 | 1493 | 95.3% | — | — | 1M | $0.75 / $3.75 |
| DeepSeek V4.1 Flash | DeepSeek | Yes (MIT) | 39 | — | 90.9%* | — | 27% | 1M | $0.15 / $0.60 |
Reading the table: For the first time since Fable 5.1 and Astra tied at 53, there's real separation at the top — Opus 5.5's 58 is a five-point lead that no other model has approached this week. Astra still holds GPQA Diamond outright at 96.1%, and its Code Arena: WebDev win nudged up to 1,800 points, ahead of Fable 5.1 Max's 1,758. The open-weight ceiling moved for the first time in weeks too: MiMo-V2.6-Pro's 46 is two points above GLM-5.3 and Kimi K3, and it landed in a flat tie with Grok 4.7 — an open model from a phone maker matching a hyped closed release, on Index score if nothing else. On price, DeepSeek V4.1 Flash is still cheapest per token at $0.15/$0.60, but MiMo-V2.6-Pro undercuts every closed model in the table by a wide margin while scoring competitively, which is the more interesting price/performance story this week.
Benchmark & Leaderboard Movement
- Opus 5.5 broke the three-way logjam at the top of the AA Index. Fable 5.1 and GPT-6 Astra had been tied at 53 for weeks; Opus 5.5 opened a five-point gap at 58, the first clear #1 since that tie formed.
- MiMo-V2.6-Pro is the new open-weight Index leader at 46, up from GLM-5.3 and Kimi K3's 44 — a real, if modest, move in the open ceiling, and it happened the same day two closed flagships launched.
- Grok 4.7's parameter count grew 40% for a two-point Index gain. The model went from 1.5T to 2.1T parameters and from an Index score of 44 to 46, while falling well behind on Terminal-Bench 4.0 (26% vs. Astra's 60% and Fable 5.1's 55%) — one of the largest hype-to-benchmark gaps this beat has tracked.
- DeepSeek V4.1 Flash quietly beat Grok 4.7 on Terminal-Bench 4.0 (27% vs. 26%) despite costing a fraction as much, a detail buried in the same Decoder piece that broke down Grok 4.7's numbers.
- The AA Index moved to v4.3.2 on September 19, a calibration-only update (re-anchoring GDPval-AA's Elo scale to DeepSeek V4.1 Flash at 1600, refined fitting for GDPval-AA and AA-Briefcase) rather than a benchmark swap — scores are comparable to last week's v4.3, unlike the Terminal-Bench 2.1-to-4.0 jump on September 7.
- Qwen 4 is officially in training, previewed in four tiers at Alibaba's Apsara Conference on September 22, with Qwen 4.5 and Qwen 5 pitched on the roadmap at 5–10 trillion parameters — none of it shipping yet.
Analysis
For agentic coding, Opus 5.5 now leads on Terminal-Bench 4.0 at 59.6% (tied with GPT-6 Astra) and claims 89.9% on SWE-bench Pro by its own harness — worth a grain of salt until someone outside Anthropic reruns it. For reasoning, GPT-6 Astra's 96.1% on GPQA Diamond remains unmatched. For cheap-and-fast, MiMo-V2.6-Pro just became the most interesting option in the field: open weights, a 1M context window, native multimodality, and pricing under a dollar per million tokens combined, all while beating the two established open-weight leaders. For open-vs-closed, the gap actually widened this week even as the open ceiling rose — Opus 5.5's jump to 58 outpaced MiMo's two-point gain to 46, so the frontier's lead grew from 9 points to 12 in the same seven days. And for anyone tempted by Grok 4.7's parameter count: bigger clearly isn't better here, not against a field where a Chinese phone company's open model just matched it on the composite score at a fraction of the price.
Sources
- Vellum — Claude Opus 5.5 Benchmarks Explained
- VentureBeat — Anthropic releases Claude Opus 5.5, beating Fable 5.1 on key agentic benchmarks at 60% cheaper API price
- Artificial Analysis — Claude Opus 5.5 model page
- The Decoder — Claude Opus 5.5 matches Fable 5.1 performance at lower cost
- TechCrunch — OpenAI launches GPT-6 Sol and Luna
- The New Stack — OpenAI releases GPT-6 Sol, Luna, and cuts token prices in half
- Artificial Analysis — GPT-6 Sol and Luna push the cost efficiency frontier
- OpenAI — GPT-6 Sol model card
- MacRumors — OpenAI's New GPT-6 Sol and Luna Models
- MarkTechPost — SpaceXAI Releases Grok 4.7
- The Decoder — xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6
- Artificial Analysis — Benchmarking Grok 4.7
- Artificial Analysis — Grok 4.7 model page
- llm-stats — Grok 4.7 Release: Same Price, Longer Horizons
- VentureBeat — Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model in the world
- TechNode — Xiaomi open-sources MiMo-V2.6 models after scaling reinforcement learning
- SiliconANGLE — Xiaomi introduces MiMo-V2.6 series open-source AI model family
- Yotta Labs — Qwen 4: Release Date, What's Confirmed, and How to Prepare
- Pandaily — Alibaba Puts Qwen4 Family Into Training; Roadmap Points to 5–10T Qwen4.5 and Qwen5
- Artificial Analysis — LLM Leaderboard
- Artificial Analysis — Terminal-Bench 4.0 Benchmark Leaderboard
- Artificial Analysis — GPQA Diamond Benchmark Leaderboard
- Artificial Analysis — Changelog
- Arena.ai — Text leaderboard
- Arena.ai — Code Arena: WebDev leaderboard
- MorphLLM — SWE-bench Pro Leaderboard (September 2026)
- BenchLM.ai — SWE-bench Pro Leaderboard: Claude Opus 5.5 Leads at 89.9%
- CodingFleet — SWE-bench Pro Leaderboard 2026
- The Register — Zuck's Muse to Spark joy with open weights release 'soon'
- My-Library — AI Model & Benchmark Watch, September 18, 2026 edition
More from News