AI Model & Benchmark Watch — October 2, 2026
Anthropic's mid-tier Sonnet 5.5 just out-scored its own flagship on the one benchmark that's supposed to matter most for coding, OpenAI refreshed Sol a week after Astra's debut, Google shipped a…
AI Model & Benchmark Watch — October 2, 2026
Anthropic's mid-tier Sonnet 5.5 just out-scored its own flagship on the one benchmark that's supposed to matter most for coding, OpenAI refreshed Sol a week after Astra's debut, Google shipped a Gemini 4 almost nobody can use yet, and not a single open-weight lab released anything new.
Overview
The closed labs had a loud week; the open-weight world had a silent one. Anthropic's Claude Sonnet 5.5 (September 28) landed two points behind Opus 5.5 on the Artificial Analysis Index — 56 versus 58 — then turned around and beat it on Artificial Analysis's own Terminal-Bench 4.0 leaderboard, 63.6% to 59.6%, while costing a fifth as much per token. OpenAI answered at DevDay with GPT-6.1 Sol (September 29), a four-point Index bump over the model it replaces at the same price. Google DeepMind shipped something it's calling Gemini 4 Argon (September 30) that debuted at #1 on the LMArena text leaderboard, but almost no one can test that claim: access is limited to vetted cybersecurity partners for now. Meanwhile Qwen 4 is still in training, GLM-5.5 is still a rumor, and no open-weight lab put out a new general-purpose model in this window — the first week this beat has tracked where that's been true.
New & Updated Models (this week)
Commercial / closed
Claude Sonnet 5.5 — Anthropic (September 28, 2026)
Who / License: Anthropic; closed, API and product access only.
What's notable: Sonnet 5.5 is the mid-tier refresh that followed Opus 5.5 by six days, and it's the more interesting release of the two. On the Artificial Analysis Index it scores 56, two points under Opus 5.5's 58 and three ahead of GPT-6 Astra's 53. But on Artificial Analysis's own Terminal-Bench 4.0 leaderboard — the independently run one, not a vendor's self-reported number — Sonnet 5.5 tops the chart at 63.6%, ahead of Opus 5.5's 59.6% and GPT-6 Astra's 58.2%. Anthropic's own benchmarking (a separate, unranked supplemental run) shows the same ordering at higher absolute numbers: Sonnet 5.5 at 70.6% against Opus 5.5's 66.4%. Either way, the cheaper model is currently winning on agentic terminal work. Pricing holds at $2/$10 per million tokens, unchanged from Sonnet 5, with a 1M-token context window. Anthropic says it's 30%+ faster than its predecessor and costs up to 30% less per task, driven by fewer tokens and tool calls rather than a lower sticker price. Claude Haiku 5.5, previewed alongside Opus 5.5 on September 22, still hasn't shipped.
Source: Anthropic — Claude Sonnet 5.5 · Anthropic — Claude Sonnet 5.5 System Card · VentureBeat — Anthropic launches Claude Sonnet 5.5 · SiliconANGLE — Anthropic debuts Claude Sonnet 5.5 · Artificial Analysis — Claude Sonnet 5.5 · Artificial Analysis — Terminal-Bench 4.0 leaderboard · CodingFleet — Terminal-Bench 4.0 Leaderboard
GPT-6.1 Sol — OpenAI (September 29, 2026, DevDay)
Who / License: OpenAI; closed, ChatGPT and API, announced at DevDay 2026 in San Francisco.
What's notable: A one-week-later refresh of GPT-6 Sol, OpenAI's mid-tier model that sits below GPT-6 Astra. It scores 52 on the AA Index, up from Sol's 48, at the same $2/$10 per-million-token price; cached input drops from $0.20 to $0.10. Context stays around 1.05M tokens. OpenAI's framing leans on cost-efficiency rather than raw capability: on Terminal-Bench Science 0.1, it scores 68.1% at $5.47 a task against Astra's $23.80, and on OSWorld 2.0 it lands within about two points of Astra at roughly a seventh of the cost. DevDay also introduced "dots," always-on background agents in ChatGPT running on Astra, and a $500/month "Pro 500" tier with access to an "Astra Ultrafast" variant — neither is a new base model, so neither gets its own entry here.
Source: OpenAI — DevDay 2026 Recap · DataCamp — GPT-6.1 Sol · Vals AI — GPT-6.1 Sol · llm-stats — GPT-6.1 Sol · Artificial Analysis — GPT-6.1 Sol
Gemini 4 Argon — Google DeepMind (September 30, 2026)
Who / License: Google DeepMind; closed, severely limited access.
What's notable: This is less a product launch than a controlled preview. Two days before it shipped, DeepMind SVP Koray Kavukcuoglu told The Information that Gemini 4 had entered "early post-training" and that Google planned a staged rollout — ship an early checkpoint, iterate on feedback — rather than wait for a polished final model. Argon looks like that first checkpoint: it's restricted to vetted cyber defenders through Google's Fairwind Program, with paid API and Google AI Ultra access promised "next" and no date attached. It scores 53 on the AA Index, tied with GPT-6 Astra and Claude Fable 5.1, and debuted at #1 on the arena.ai text leaderboard at 1525 Elo — ahead of every Claude Opus or Fable variant currently ranked. The headline spec is a 1M-token output ceiling, up from prior Gemini generations' 64K, aimed at long, multi-step agentic work; Google's own examples include catching a vulnerability in hospital software that "earlier frontier models missed" and a Rust video decoder optimization running 2.7x faster. Text and image input, text output, 1M total context, priced at $2/$10 per million tokens on Artificial Analysis's listing.
Source: gHacks — Google launches Gemini 4 Argon with a 1 million token output limit · Dataconomy — DeepMind says Gemini 4 is coming much earlier than expected · The Information — Google nears release of flagship Gemini 4 AI model · Artificial Analysis — Gemini 4 Argon · Arena.ai — Text leaderboard
Also this week: Cohere released Embed 5 (October 1), a closed embedding model in Pro and Fast tiers — not a chat/reasoning model, so no table entry, but it claims to beat Voyage 4 Large and Gemini Embedding 2 on ViDoRe V3 (85.8 and 84.5 versus 83.7 and 83.2). Priced at $0.08–$0.12 per million tokens.
Source: Cohere — Embed 5
Open-weight
Nothing shipped in the general-purpose chat/reasoning category this week — a first for this beat. Qwen 4 remains in training after its September 22 Apsara preview, with no date, weights, or benchmarks. GLM-5.5 is still rumored, not released. DeepSeek, Kimi, and Meta's open-weight plans had no news in this window.
The one genuine open-weight release was narrower: Cloudflare put out Clef and Clef-flash (October 1), Apache-2.0 "decision models" built on Qwen3.8 bases (27B and 9B) that return calibrated probabilities over a fixed set of answers rather than generating free text — useful for classification and routing, not a general chat model. Cloudflare claims a 94.20 macro-F1 on BANKING77 against a comparison baseline of 79.74, and median latency of 38.8ms for Clef-flash. Worth knowing about if you're doing high-volume classification, not a frontier contender.
Source: Cloudflare — Clef decision models · Hugging Face — Cloudflare/clef
Head-to-Head — Current Frontier
State of play as of October 2, 2026. AA Index stays on v4.3.2, unchanged since the September 19 calibration update — scores remain comparable to last week. Terminal-Bench 4.0 figures are Artificial Analysis's own independently run public-board scores where available; a few labs (Anthropic among them) publish separate, higher "supplemental" numbers from their own harnesses that Artificial Analysis itself says shouldn't be read as like-for-like with the public board — those vendor numbers are called out in the model sections above, not the table. LMArena Elo is arena.ai's text leaderboard; new releases take time to accumulate votes, so several of this week's launches show "—." GPQA Diamond and SWE-bench cells are marked "—" where no current, source-backed figure could be confirmed — several labs (Google, xAI, Xiaomi) simply don't publish one.
| Model | Org | Open? | AA Index (v4.3.2) | LMArena Elo | GPQA Diamond | SWE-bench (Verified/Pro) | Terminal-Bench 4.0 | Context | $ / Mtok in/out |
|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 | Anthropic | No | 56 | — (too new to rank) | — | — | 63.6% | 1M | $2 / $10 |
| Claude Opus 5.5 | Anthropic | No | 58 | 1504 (high) | — | 89.9% (Pro)* | 59.6% | 1M | $4 / $20 |
| Claude Fable 5.1 | Anthropic | No | 53 | 1501 (max) | 92.6% | 81.2% (Pro) | 57.9% | 1M | $10 / $50 |
| GPT-6 Astra | OpenAI | No | 53 | — (not in top ranks) | 96.1% | — | 58.2% | 1M | $10 / $50 |
| Gemini 4 Argon | Google DeepMind | No | 53 | 1525 (high, gated access) | — | — | — | 1M | $2 / $10 |
| GPT-6.1 Sol | OpenAI | No | 52 | — | — | — | — | 1.05M | $2 / $10 |
| Muse Spark 1.3 | Meta | No (weights undecided) | 48 | 1495 (max) | — | — | — | 1M | $1.25 / $4.25 |
| Grok 4.7 | xAI | No | 46 | — | — | — | 37.6% | 500K | $2 / $6 (<200K) |
| MiMo-V2.6-Pro | Xiaomi | Yes (MIT) | 46 | — | — | — | — | 1M | $0.43 / $0.87 |
| GLM-5.3 | Zhipu / Z.ai | Yes (custom license, MaaS clause) | 45 | — | — | 77.8% (Verified) | 41.8% | 1M | $1.40 / $4.40 |
| Kimi K3 | Moonshot AI | Yes (custom license) | 44 | — | 93.5%* | 93.40% (Verified) | — | 1.05M | $3 / $15 |
| Gemini 3.8 Flash | Google DeepMind | No | 41 | 1494 (high) | 95.3% | — | 19.7%† | 1M | $0.75 / $3.75 |
| DeepSeek V4.1 Flash | DeepSeek | Yes (MIT) | 39 | — | 90.9%* | — | 27%* | 1M | $0.30 / $1.20 |
* Vendor- or launch-reported, not independently reproduced on the current benchmark version. † Scored under the September 7 Terminal-Bench 4.0 hardening pass that dropped several models sharply from their 2.1-era numbers — a methodology change, not new news this week.
Reading the table: The top of the Index is now a four-way knot — Opus 5.5 at 58, with Fable 5.1, GPT-6 Astra, and Gemini 4 Argon all tied at 53, and Sonnet 5.5 sitting between them at 56. But Index rank and Terminal-Bench rank have split: Sonnet 5.5 leads the agentic-coding board outright, ahead of both Opus 5.5 and Astra, despite being the cheapest model in that top tier. GPT-6 Astra still owns GPQA Diamond at 96.1%, untouched by anyone this week. Among open models, GLM-5.3's 41.8% on Terminal-Bench 4.0 is the best verified open-weight score on that eval — ahead of Grok 4.7's 37.6%, a closed model xAI shipped eleven days earlier. MiMo-V2.6-Pro still holds the open-weight Index lead at 46, but nothing moved there this week; the open ceiling is exactly where it was on September 25.
Benchmark & Leaderboard Movement
- Sonnet 5.5 beats its own flagship on Terminal-Bench 4.0. 63.6% versus Opus 5.5's 59.6% and GPT-6 Astra's 58.2%, on Artificial Analysis's independently run public board — the cheaper, faster mid-tier model is currently the one to reach for on agentic terminal work, not the $4/$20 flagship.
- Gemini 4 Argon debuted at #1 on arena.ai's text leaderboard (1525 Elo) and tied for the top AA Index slot among this week's three closed launches at 53 — but almost nobody can verify either claim firsthand, since access is restricted to Google's Fairwind cyber-defense partners.
- An open-weight model beat a hyped closed one on Terminal-Bench 4.0 for the second week running. Last week it was DeepSeek V4.1 Flash edging Grok 4.7 (27% vs. 26%); this week GLM-5.3 posted 41.8% against Grok 4.7's current 37.6%.
- No open-weight general-purpose model shipped this week — Qwen 4 still in training, GLM-5.5 still a rumor, DeepSeek and Kimi quiet. The open-weight ceiling (MiMo-V2.6-Pro, AA Index 46) hasn't moved since September 25.
- GPT-6 Sol had a one-week shelf life. OpenAI replaced it with GPT-6.1 Sol at DevDay, a four-point Index gain at an unchanged price — one of the fastest full-cycle refreshes either major lab has run this year.
Analysis
For agentic coding, the model to reach for just changed: Sonnet 5.5 leads Terminal-Bench 4.0 outright and costs a fifth of Opus 5.5 per token, though Anthropic still positions Opus for "complex, open-ended work requiring sustained judgment" where Sonnet hasn't been tested as thoroughly. For reasoning, GPT-6 Astra's GPQA Diamond lead (96.1%) is unchallenged. For cheap-and-fast, nothing beats MiMo-V2.6-Pro or DeepSeek V4.1 Flash on price per token, and neither of this week's new models even competes in that tier. For long-context agentic work, Gemini 4 Argon's 1M-token output ceiling is a real capability jump if you can get access — which, for now, you can't. On open versus closed, the gap held roughly steady on the Index (best open still 46, best closed now 58) even as it narrowed on specific evals: an open model beat a closed flagship-adjacent release on Terminal-Bench 4.0 for the second straight week, which is a more interesting signal than the composite score gap suggests.
Sources
- Anthropic — Claude Sonnet 5.5
- Anthropic — Claude Sonnet 5.5 System Card
- VentureBeat — Anthropic launches Claude Sonnet 5.5 with 30% cost reduction per task
- SiliconANGLE — Anthropic debuts Claude Sonnet 5.5, running 30% faster
- OpenAI — DevDay 2026 Recap
- DataCamp — GPT-6.1 Sol
- Vals AI — GPT-6.1 Sol
- llm-stats — GPT-6.1 Sol
- gHacks — Google launches Gemini 4 Argon with a 1 million token output limit
- Dataconomy — DeepMind says Gemini 4 is coming much earlier than expected
- The Information — Google nears release of flagship Gemini 4 AI model
- Cohere — Embed 5
- Cloudflare — Clef decision models
- Hugging Face — Cloudflare/clef
- Pandaily — Alibaba puts Qwen4 family into training
- CellCog — Qwen 4 release date
- Artificial Analysis — LLM Leaderboard
- Artificial Analysis — Claude Sonnet 5.5
- Artificial Analysis — Claude Opus 5.5
- Artificial Analysis — Claude Fable 5.1
- Artificial Analysis — GPT-6 Astra
- Artificial Analysis — GPT-6.1 Sol
- Artificial Analysis — Gemini 4 Argon
- Artificial Analysis — Gemini 3.8 Flash
- Artificial Analysis — Muse Spark 1.3
- Artificial Analysis — MiMo-V2.6-Pro
- Artificial Analysis — GLM-5.3
- Artificial Analysis — Kimi K3
- Artificial Analysis — DeepSeek V4.1 Flash
- Artificial Analysis — Terminal-Bench 4.0 Benchmark Leaderboard
- Artificial Analysis — GPQA Diamond Benchmark Leaderboard
- Artificial Analysis — Changelog
- CodingFleet — Terminal-Bench 4.0 Leaderboard: Grok 4.7, GPT-6 & New Runs
- Flowtivity — Terminal-Bench 4.0 Exposes Inflated AI Scores
- Arena.ai — Text leaderboard
- My-Library — AI Model & Benchmark Watch, September 25, 2026 edition
More from News