← September 2026
News 2026-09-18

AI Model & Benchmark Watch — September 18, 2026

Grok 4.7 missed a third deadline and got graded by its own creator before anyone outside xAI could touch it, while Artificial Analysis rewrote the coding benchmark under the whole leaderboard for the…

AI Model & Benchmark Watch — September 18, 2026

AI Model & Benchmark Watch — September 18, 2026

Grok 4.7 missed a third deadline and got graded by its own creator before anyone outside xAI could touch it, while Artificial Analysis rewrote the coding benchmark under the whole leaderboard for the second time in two weeks.

Overview

No flagship closed model shipped this week. Grok 4.7 was supposed to land September 12; instead Elon Musk posted that it "needs a few more days to cook," then on September 14 rated his own unreleased model as "roughly on par with Opus 5.0, not 5.1" — a strange thing to know about a model nobody outside xAI has run. The real story is procedural again: Artificial Analysis pushed its Intelligence Index to v4.3 on September 7, upgrading Terminal-Bench from 2.1 to 4.0 and swapping a banking-agent eval for a private set of Zapier workflow tasks. The composite scores barely moved, but the Terminal-Bench column collapsed across the board — GPT-6 Astra went from 87% to 58% on paper without changing at all. Two smaller launches filled the gap: Sakana AI shipped a pair of orchestration models that route work across a pool of other models rather than being models themselves, and Shanghai AI Lab quietly open-sourced Atria Dawn Preview, a 744B-parameter research-agent model built on GLM-5.2.

New & Updated Models (this week)

Commercial / closed

Nothing from the usual flagship labs. Grok 4.7 is now on its third missed date — Musk's September 2 announcement targeted September 12, that came and gone, and as of this writing xAI has published no model card, API ID, price, or benchmark table. Musk's own September 14 assessment, posted before release, put it "roughly on par with Opus 5.0, not 5.1... better in some ways, worse in others."

Source: iweaver — Grok 4.7 Release Date, Features & Latest Updates · techjournal.org — Grok 4.7 Release Date: Why xAI Delayed It to September · BigGo Finance — Musk Announces Grok 4.7 Launch in Ten Days, Touts 2.1 Trillion Parameters

Fugu Max v1.0 and Fugu Ultra v2.0 — Sakana AI (September 11, 2026)

Who / License: Sakana AI; closed API only, no published weights.

What's notable: These aren't foundation models in the usual sense — they're orchestrators. A router model takes your request, decides which models in a fixed pool (including Nvidia's Nemotron family) should handle which piece of it, and can call instances of itself recursively before stitching the answers back together. Which models are in the pool and how they're picked stays undisclosed. Fugu Max runs $2/$6 per million tokens; Fugu Ultra v2.0 has a 1M-token context window at $5/$30 and, on Sakana's own Chartography visual-reasoning benchmark, scored 48.3 against Claude Opus 5's 27.3 and Claude Fable 5's 29.5. Both are drop-in replacements for the original Fugu via a single API parameter change.

Source: Sakana AI — Introducing Fugu Max and Fugu Ultra v2 · DataNorth — Sakana AI Launches Fugu Max and Fugu Ultra v2 · MarkTechPost — Sakana AI Launches Fugu Max and Fugu Ultra v2

Open-weight

Atria Dawn Preview — Shanghai AI Lab / InternLM (September 11–15, 2026)

Who / License: Released under the Atria name by Shanghai AI Lab; open weights on Hugging Face under MIT.

What's notable: A 744B-parameter mixture-of-experts model built on the GLM-5.2 foundation, aimed squarely at long-horizon research agents — the pitch is carrying a method from a paper through to executable experiments and a report someone else can check. It has a 256K context window and an explicit reasoning mode that trades latency for accuracy. DeepSeek-style caveats apply: the benchmark table is vendor-reported, including a 59.6% on SWE-bench Pro, and none of it has been independently reproduced yet.

Source: The Globe and Mail — ATRIA Releases Atria Dawn Preview for Long Horizon Research Agents · emergent.sh — Shanghai AI Lab Launches Atria Dawn Preview Model · Hugging Face — internlm/Atria-Dawn-Preview · OrcaRouter — Atria Dawn Preview: InternLM's 744B Agentic MoE Ships Quietly


Head-to-Head — Current Frontier

State of play as of September 18, 2026. AA Index uses Artificial Analysis's v4.3, which launched September 7 and upgraded Terminal-Bench from 2.1 to 4.0 while swapping τ³-Banking for AutomationBench-AA, a private set of 657 SaaS workflow tasks. The composite scores held nearly steady versus last week's v4.2 (a handful of points moved for individual models), but Terminal-Bench 4.0 numbers are not comparable to the Terminal-Bench 2.1 figures printed in earlier editions — the new version uses harder tasks, so a lower score here does not mean a model got worse. SWE-bench Verified stays frozen at the values vals.ai archived on May 4, 2026 as saturated and contaminated; the Pro variant, run independently by third parties, is still moving. LMArena Elo is Arena.ai's text leaderboard; "—" means the model hasn't accumulated enough votes to rank on that specific board, is measured on a different arena (like Astra's coding-specific win, noted below), or isn't yet listed. Figures marked * are self-reported by the vendor on its own harness.

Model Org Open? AA Index (v4.3) LMArena Elo GPQA Diamond SWE-bench (Verified/Pro) Terminal-Bench 4.0 Context $ / Mtok in/out
Claude Fable 5.1 Anthropic No 53 1498 92.6% 81.2% (Pro) 57.9% 1M $10 / $50
GPT-6 Astra OpenAI No 53 — (Code Arena #1) 96.1% — 58.2% 1.05M $10 / $50
Claude Opus 5 Anthropic No 51 1493 84.1% 97.00% (Verified, frozen) / 79.2% (Pro) 51.8% 1M $5 / $25
Muse Spark 1.3 Meta No (weights undecided) 48 1493 — — — 1M $1.25 / $4.25
GPT-5.6 Sol OpenAI No 47 — 94.1% 64.6% (Pro)* — 1M+ $5 / $30
GLM-5.3 Zhipu / Z.ai Yes (custom license, MaaS clause) 44 — — — 41.8% 1M $1.40 / $4.40
Kimi K3 Moonshot AI Yes (custom license) 44 — 93.5% 93.40% (Verified) — 1.05M $3 / $15
Grok 4.6 xAI No 44 — 94.9% — — 500K $2 / $6 (below 200K)
Gemini 3.8 Flash Google DeepMind No 41 1493 95.3% — — 1M $0.75 / $3.75
Qwen3.8-Max Alibaba No 40 — 92.6%* 67.7% (Pro)* — 1M $2 / $6
DeepSeek V4.1 Flash DeepSeek Yes (MIT) 40 — 90.9%* — — 1M $0.15 / $0.60
Atria Dawn Preview Shanghai AI Lab Yes (MIT) — — — 59.6% (Pro)* — 256K —

Reading the table: Fable 5.1 and Astra are still tied atop the Index at 53, and the gap between them keeps splitting along the same lines — Astra owns GPQA Diamond (96.1%) and now Terminal-Bench 4.0 (58.2%) on independently-run numbers, while Fable 5.1 leads the composite and, oddly, isn't even Anthropic's best performer on Arena's text board: Claude Fable 5, a model one version behind, sits at 1506 versus Fable 5.1's 1498. Astra's actual head-to-head win this week came on a different board entirely — Arena's Code Arena: WebDev, where it took #1 at 1797 points, 35 ahead of Fable 5.1's 1762 and well clear of Opus 5's 1688. On price, DeepSeek V4.1 Flash is still the standout on paper at $0.15/$0.60 per million tokens, tied with Qwen3.8-Max on the Index — but see below, because its coding claims took a hit this week.


Benchmark & Leaderboard Movement

  • Artificial Analysis moved to Index v4.3 on September 7, upgrading Terminal-Bench to version 4.0 (harder tasks, tighter grading) and replacing τ³-Banking with AutomationBench-AA, a private 657-task benchmark built from real Zapier workflows across finance, HR, marketing, ops, sales, and support. Composite scores barely shifted — Fable 5.1 and Astra still tied at 53 — but every model's Terminal-Bench number dropped 25-plus points on paper (Astra: 87.3% → 58.2%; Fable 5.1: 85.0% → 57.9%; Opus 5: 84.6% → 51.8%) purely because the ruler changed.
  • GPT-6 Astra's independent Code Arena win is new and real. It took #1 on Arena's Code Arena: WebDev leaderboard at 1797 points, a genuine 35-point lead over Fable 5.1 — a different signal than the general text Arena, where Astra hasn't cracked the top ranks at all.
  • Arena backfilled votes from its "Battles in Direct" experiment, which has been running since March and started counting toward leaderboard scores in June. The backfill nudges scores for models active during that window; Fable 5.1's Elo moved from 1504 last edition to 1498 this week, a methodology shift rather than a capability change.
  • DeepSeek's DeepSWE win over Opus 5 doesn't survive independent reproduction. Datacurve's own DeepSWE v1.1 leaderboard places Muse Spark 1.3 first at 75.4, ahead of DeepSeek V4.1 Flash's 74.2 — the opposite of the ordering DeepSeek published on its own harness last week.
  • Grok 4.7 slipped a third deadline. The September 12 target Musk set September 2 came and went with no model card, and Musk himself pre-graded the model against Opus 5.0 on September 14, before anyone outside xAI had touched it.

Analysis

For agentic coding, the field just got harder to compare, not easier — Terminal-Bench's jump to version 4.0 reset the scale for everyone, and only Astra, Fable 5.1, Opus 5, and GLM-5.3 have confirmed numbers on it so far. Astra leads that fresh board at 58.2% and also owns the independently-run Code Arena win; Opus 5 still holds the frozen SWE-bench Verified record from May. For reasoning, Astra's GPQA Diamond score (96.1%) remains the most solid number in the table. For cheap-and-fast, DeepSeek V4.1 Flash still posts the best price on paper, but this week's Datacurve reproduction is a reminder to verify its agentic claims before trusting them over Muse Spark 1.3 or Qwen3.8-Max. For open-weight, Atria Dawn Preview is worth watching once someone outside Shanghai AI Lab runs its numbers, but today's open ceiling is still GLM-5.3 and Kimi K3 at an Index of 44, nine points behind the closed frontier — unchanged from last week.

The recurring theme across three straight editions now: benchmarks keep getting rebuilt out from under the models being scored. GPQA Diamond got demoted for saturation on September 4. SWE-bench Verified has been frozen since May. Terminal-Bench jumped a full version on September 7. None of that reflects any model changing — it reflects trackers responding to models getting good enough to break the old test. Worth remembering the next time a headline score looks like it moved.


Sources

More from News