AI Model & Benchmark Watch — September 18, 2026
Grok 4.7 missed a third deadline and got graded by its own creator before anyone outside xAI could touch it, while Artificial Analysis rewrote the coding benchmark under the whole leaderboard for the…
AI Model & Benchmark Watch — September 18, 2026
Grok 4.7 missed a third deadline and got graded by its own creator before anyone outside xAI could touch it, while Artificial Analysis rewrote the coding benchmark under the whole leaderboard for the second time in two weeks.
Overview
No flagship closed model shipped this week. Grok 4.7 was supposed to land September 12; instead Elon Musk posted that it "needs a few more days to cook," then on September 14 rated his own unreleased model as "roughly on par with Opus 5.0, not 5.1" — a strange thing to know about a model nobody outside xAI has run. The real story is procedural again: Artificial Analysis pushed its Intelligence Index to v4.3 on September 7, upgrading Terminal-Bench from 2.1 to 4.0 and swapping a banking-agent eval for a private set of Zapier workflow tasks. The composite scores barely moved, but the Terminal-Bench column collapsed across the board — GPT-6 Astra went from 87% to 58% on paper without changing at all. Two smaller launches filled the gap: Sakana AI shipped a pair of orchestration models that route work across a pool of other models rather than being models themselves, and Shanghai AI Lab quietly open-sourced Atria Dawn Preview, a 744B-parameter research-agent model built on GLM-5.2.
New & Updated Models (this week)
Commercial / closed
Nothing from the usual flagship labs. Grok 4.7 is now on its third missed date — Musk's September 2 announcement targeted September 12, that came and gone, and as of this writing xAI has published no model card, API ID, price, or benchmark table. Musk's own September 14 assessment, posted before release, put it "roughly on par with Opus 5.0, not 5.1... better in some ways, worse in others."
Source: iweaver — Grok 4.7 Release Date, Features & Latest Updates · techjournal.org — Grok 4.7 Release Date: Why xAI Delayed It to September · BigGo Finance — Musk Announces Grok 4.7 Launch in Ten Days, Touts 2.1 Trillion Parameters
Fugu Max v1.0 and Fugu Ultra v2.0 — Sakana AI (September 11, 2026)
Who / License: Sakana AI; closed API only, no published weights.
What's notable: These aren't foundation models in the usual sense — they're orchestrators. A router model takes your request, decides which models in a fixed pool (including Nvidia's Nemotron family) should handle which piece of it, and can call instances of itself recursively before stitching the answers back together. Which models are in the pool and how they're picked stays undisclosed. Fugu Max runs $2/$6 per million tokens; Fugu Ultra v2.0 has a 1M-token context window at $5/$30 and, on Sakana's own Chartography visual-reasoning benchmark, scored 48.3 against Claude Opus 5's 27.3 and Claude Fable 5's 29.5. Both are drop-in replacements for the original Fugu via a single API parameter change.
Source: Sakana AI — Introducing Fugu Max and Fugu Ultra v2 · DataNorth — Sakana AI Launches Fugu Max and Fugu Ultra v2 · MarkTechPost — Sakana AI Launches Fugu Max and Fugu Ultra v2
Open-weight
Atria Dawn Preview — Shanghai AI Lab / InternLM (September 11–15, 2026)
Who / License: Released under the Atria name by Shanghai AI Lab; open weights on Hugging Face under MIT.
What's notable: A 744B-parameter mixture-of-experts model built on the GLM-5.2 foundation, aimed squarely at long-horizon research agents — the pitch is carrying a method from a paper through to executable experiments and a report someone else can check. It has a 256K context window and an explicit reasoning mode that trades latency for accuracy. DeepSeek-style caveats apply: the benchmark table is vendor-reported, including a 59.6% on SWE-bench Pro, and none of it has been independently reproduced yet.
Source: The Globe and Mail — ATRIA Releases Atria Dawn Preview for Long Horizon Research Agents · emergent.sh — Shanghai AI Lab Launches Atria Dawn Preview Model · Hugging Face — internlm/Atria-Dawn-Preview · OrcaRouter — Atria Dawn Preview: InternLM's 744B Agentic MoE Ships Quietly
Head-to-Head — Current Frontier
State of play as of September 18, 2026. AA Index uses Artificial Analysis's v4.3, which launched September 7 and upgraded Terminal-Bench from 2.1 to 4.0 while swapping τ³-Banking for AutomationBench-AA, a private set of 657 SaaS workflow tasks. The composite scores held nearly steady versus last week's v4.2 (a handful of points moved for individual models), but Terminal-Bench 4.0 numbers are not comparable to the Terminal-Bench 2.1 figures printed in earlier editions — the new version uses harder tasks, so a lower score here does not mean a model got worse. SWE-bench Verified stays frozen at the values vals.ai archived on May 4, 2026 as saturated and contaminated; the Pro variant, run independently by third parties, is still moving. LMArena Elo is Arena.ai's text leaderboard; "—" means the model hasn't accumulated enough votes to rank on that specific board, is measured on a different arena (like Astra's coding-specific win, noted below), or isn't yet listed. Figures marked * are self-reported by the vendor on its own harness.
| Model | Org | Open? | AA Index (v4.3) | LMArena Elo | GPQA Diamond | SWE-bench (Verified/Pro) | Terminal-Bench 4.0 | Context | $ / Mtok in/out |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | No | 53 | 1498 | 92.6% | 81.2% (Pro) | 57.9% | 1M | $10 / $50 |
| GPT-6 Astra | OpenAI | No | 53 | — (Code Arena #1) | 96.1% | — | 58.2% | 1.05M | $10 / $50 |
| Claude Opus 5 | Anthropic | No | 51 | 1493 | 84.1% | 97.00% (Verified, frozen) / 79.2% (Pro) | 51.8% | 1M | $5 / $25 |
| Muse Spark 1.3 | Meta | No (weights undecided) | 48 | 1493 | — | — | — | 1M | $1.25 / $4.25 |
| GPT-5.6 Sol | OpenAI | No | 47 | — | 94.1% | 64.6% (Pro)* | — | 1M+ | $5 / $30 |
| GLM-5.3 | Zhipu / Z.ai | Yes (custom license, MaaS clause) | 44 | — | — | — | 41.8% | 1M | $1.40 / $4.40 |
| Kimi K3 | Moonshot AI | Yes (custom license) | 44 | — | 93.5% | 93.40% (Verified) | — | 1.05M | $3 / $15 |
| Grok 4.6 | xAI | No | 44 | — | 94.9% | — | — | 500K | $2 / $6 (below 200K) |
| Gemini 3.8 Flash | Google DeepMind | No | 41 | 1493 | 95.3% | — | — | 1M | $0.75 / $3.75 |
| Qwen3.8-Max | Alibaba | No | 40 | — | 92.6%* | 67.7% (Pro)* | — | 1M | $2 / $6 |
| DeepSeek V4.1 Flash | DeepSeek | Yes (MIT) | 40 | — | 90.9%* | — | — | 1M | $0.15 / $0.60 |
| Atria Dawn Preview | Shanghai AI Lab | Yes (MIT) | — | — | — | 59.6% (Pro)* | — | 256K | — |
Reading the table: Fable 5.1 and Astra are still tied atop the Index at 53, and the gap between them keeps splitting along the same lines — Astra owns GPQA Diamond (96.1%) and now Terminal-Bench 4.0 (58.2%) on independently-run numbers, while Fable 5.1 leads the composite and, oddly, isn't even Anthropic's best performer on Arena's text board: Claude Fable 5, a model one version behind, sits at 1506 versus Fable 5.1's 1498. Astra's actual head-to-head win this week came on a different board entirely — Arena's Code Arena: WebDev, where it took #1 at 1797 points, 35 ahead of Fable 5.1's 1762 and well clear of Opus 5's 1688. On price, DeepSeek V4.1 Flash is still the standout on paper at $0.15/$0.60 per million tokens, tied with Qwen3.8-Max on the Index — but see below, because its coding claims took a hit this week.
Benchmark & Leaderboard Movement
- Artificial Analysis moved to Index v4.3 on September 7, upgrading Terminal-Bench to version 4.0 (harder tasks, tighter grading) and replacing τ³-Banking with AutomationBench-AA, a private 657-task benchmark built from real Zapier workflows across finance, HR, marketing, ops, sales, and support. Composite scores barely shifted — Fable 5.1 and Astra still tied at 53 — but every model's Terminal-Bench number dropped 25-plus points on paper (Astra: 87.3% → 58.2%; Fable 5.1: 85.0% → 57.9%; Opus 5: 84.6% → 51.8%) purely because the ruler changed.
- GPT-6 Astra's independent Code Arena win is new and real. It took #1 on Arena's Code Arena: WebDev leaderboard at 1797 points, a genuine 35-point lead over Fable 5.1 — a different signal than the general text Arena, where Astra hasn't cracked the top ranks at all.
- Arena backfilled votes from its "Battles in Direct" experiment, which has been running since March and started counting toward leaderboard scores in June. The backfill nudges scores for models active during that window; Fable 5.1's Elo moved from 1504 last edition to 1498 this week, a methodology shift rather than a capability change.
- DeepSeek's DeepSWE win over Opus 5 doesn't survive independent reproduction. Datacurve's own DeepSWE v1.1 leaderboard places Muse Spark 1.3 first at 75.4, ahead of DeepSeek V4.1 Flash's 74.2 — the opposite of the ordering DeepSeek published on its own harness last week.
- Grok 4.7 slipped a third deadline. The September 12 target Musk set September 2 came and went with no model card, and Musk himself pre-graded the model against Opus 5.0 on September 14, before anyone outside xAI had touched it.
Analysis
For agentic coding, the field just got harder to compare, not easier — Terminal-Bench's jump to version 4.0 reset the scale for everyone, and only Astra, Fable 5.1, Opus 5, and GLM-5.3 have confirmed numbers on it so far. Astra leads that fresh board at 58.2% and also owns the independently-run Code Arena win; Opus 5 still holds the frozen SWE-bench Verified record from May. For reasoning, Astra's GPQA Diamond score (96.1%) remains the most solid number in the table. For cheap-and-fast, DeepSeek V4.1 Flash still posts the best price on paper, but this week's Datacurve reproduction is a reminder to verify its agentic claims before trusting them over Muse Spark 1.3 or Qwen3.8-Max. For open-weight, Atria Dawn Preview is worth watching once someone outside Shanghai AI Lab runs its numbers, but today's open ceiling is still GLM-5.3 and Kimi K3 at an Index of 44, nine points behind the closed frontier — unchanged from last week.
The recurring theme across three straight editions now: benchmarks keep getting rebuilt out from under the models being scored. GPQA Diamond got demoted for saturation on September 4. SWE-bench Verified has been frozen since May. Terminal-Bench jumped a full version on September 7. None of that reflects any model changing — it reflects trackers responding to models getting good enough to break the old test. Worth remembering the next time a headline score looks like it moved.
Sources
- iweaver — Grok 4.7 Release Date, Features & Latest Updates
- techjournal.org — Grok 4.7 Release Date: Why xAI Delayed It to September
- BigGo Finance — Musk Announces Grok 4.7 Launch in Ten Days, Touts 2.1 Trillion Parameters to Beat All Models
- explainx.ai — Grok 4.7 Release Date and Specs: 2.1T Params
- LatestLY — Elon Musk Confirms Grok 4.7 Release in 10 Days
- Sakana AI — Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier
- DataNorth — Sakana AI Launches Fugu Max and Fugu Ultra v2
- MarkTechPost — Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration
- AiCybr — Sakana Fugu Max and Ultra v2: Pricing, Benchmarks, API and Orchestration Architecture
- theroboticsmedia — Sakana AI Ships Fugu Ultra v2.0, a 1M-Token Multi-Agent Orchestration Model
- OpenRouter — Fugu Ultra v2
- The Globe and Mail — ATRIA Releases Atria Dawn Preview for Long Horizon Research Agents
- emergent.sh — Shanghai AI Lab Launches Atria Dawn Preview Model
- Hugging Face — internlm/Atria-Dawn-Preview
- OrcaRouter — Atria Dawn Preview: InternLM's 744B Agentic MoE Ships Quietly
- OrcaRouter — Atria Dawn vs GPT-5.6 Sol: The 13-Point Index Gap, Re-Run
- HokAI — Atria Dawn Preview: 256K Context, 59.6% SWE-bench
- Artificial Analysis — Announcing the Artificial Analysis Intelligence Index v4.3
- Artificial Analysis — LLM Leaderboard
- Artificial Analysis — Terminal-Bench 4.0 Benchmark Leaderboard
- CodingFleet — Terminal-Bench 4.0 Leaderboard 2026: AI Agents Ranked by CLI Work
- Snorkel AI — Terminal-Bench 4.0
- Arena.ai — Text leaderboard
- Arena.ai — Leaderboard Changelog
- Arena.ai on X — GPT-6 Astra takes #1 on Code Arena: WebDev
- CryptoBriefing — OpenAI's GPT-6 Astra tops Code Arena: WebDev with 1,797 points
- llm-stats.com — GPT-6 Astra Benchmarks, Pricing & Context Window
- OpenRouter — GPQA Diamond Leaderboard
- IntuitionLabs — GPQA Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
- MorphLLM — SWE-bench Pro Leaderboard (September 2026)
- MorphLLM — Claude Benchmarks (2026): Opus 5, Sonnet 5, and Fable 5 at 95% SWE-bench Verified
- CodingFleet — SWE-bench Pro Leaderboard: Claude Fable 5.1 Takes #1 at 81.2%
- Digital Applied — DeepSeek V4.1 Flash: Benchmarks, Prices and Pro Cutoff
- MindStudio — DeepSeek V4.1 Flash Benchmarks vs Opus 5 and GPT-5.6: What's Real?
- OrcaRouter — DeepSeek V4.1 Flash Benchmarks: What 74.2 Really Proves
- VentureBeat — Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can't broadly use yet
- The Register — Zuck's Muse to Spark joy with open weights release 'soon'
- TheNextWeb — Meta's new AI model edges closer to OpenAI and Anthropic, its AI Chief says
- My-Library — AI Model & Benchmark Watch, September 11, 2026 edition
More from News