← September 2026
News 2026-09-04

AI Model & Benchmark Watch — September 4, 2026

Four flagship-tier models shipped in three days — Claude Fable 5.1, Meta's Muse Spark 1.3, Gemini 3.8 Flash, and OpenAI's GPT-6 Astra — after a summer where the loudest news was usually about what…

AI Model & Benchmark Watch — September 4, 2026

AI Model & Benchmark Watch — September 4, 2026

Four flagship-tier models shipped in three days — Claude Fable 5.1, Meta's Muse Spark 1.3, Gemini 3.8 Flash, and OpenAI's GPT-6 Astra — after a summer where the loudest news was usually about what hadn't shipped yet.

Overview

This is the busiest week this beat has covered since it started. Anthropic, Meta, Google, and OpenAI all put out new models between September 1 and September 3, and GLM-5.3's full open weights — missing as of last week's edition — turned up too, just under a license that quietly dropped the MIT terms Zhipu had used for every GLM release before it. Claude Fable 5.1 took the Artificial Analysis Intelligence Index lead within a day of launch. GPT-6 Astra arrived carrying the label OpenAI had spent a month warning about: the first model to cross the "Critical" threshold for cyber capability under its own Preparedness Framework, released anyway, gated behind OpenAI's Daybreak cybersecurity program at launch. Grok 4.7, still the one flagship missing from this picture, slipped again — Musk now says September 11 or 12.

New & Updated Models (this week)

Commercial / closed

Claude Fable 5.1 and Claude Mythos 5.1 — Anthropic (September 1, 2026)

Who / License: Anthropic; closed, API-only, available directly and through AWS, Google Cloud, and Azure.

What's notable: An efficiency pass on June's Fable 5 and Mythos 5 rather than a ground-up rebuild — Anthropic says it delivers similar or better results at low-to-medium reasoning effort for 25% less compute, with cached-input reads down 75%. Sticker pricing is unchanged at $10/$50 per million tokens. Fable 5.1 scores 92.6% on GPQA Diamond and 81.2% on SWE-bench Pro; Mythos 5.1, the version without Fable's extra safeguards, beats it on Terminal-Bench 4.0 (60.9% vs. 55.8%), which tells you the safety layer still costs something on agentic tasks. Notably, Anthropic's own launch materials don't lead with a SWE-bench number at all — the emphasis has shifted to Terminal-Bench-Science and long-running agentic work. 1M-token context, multimodal input, invisible watermarking on generated text with a detection API now in private preview for EU compliance.

Source: VentureBeat — Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for cache reads · TechCrunch — Anthropic's new Fable release is cheaper, less restrictive · Anthropic — Claude Fable 5.1 & Claude Mythos 5.1 System Card

GPT-6 Astra — OpenAI (September 3, 2026)

Who / License: OpenAI; closed, limited preview through the Daybreak cybersecurity program, wider ChatGPT/API/AWS rollout "in the coming days."

What's notable: OpenAI had already told Axios in early August it couldn't rule out critical cyber risk from its next model; Astra is that model, and it's the first OpenAI has classified as meeting the "Critical" cybersecurity capability threshold under its Preparedness Framework. It shipped anyway, with tighter access controls at launch. 1.05M-token context, 128K max output, multimodal input, priced at $10/$50 per million tokens ($1 cached). OpenAI's self-reported numbers: 96.0% on GPQA, 74.1% on DeepSWE v1.1, 57.7% on Terminal-Bench 4.0, 100% on ExploitBench, 99.2% on SRE-Bench reverse engineering. Independent numbers tell a flatter story — Artificial Analysis puts its Intelligence Index at 61, tied with GPT-5.6 Sol and below Fable 5.1 (66), Opus 5 (63), and Muse Spark 1.3 (62). OpenAI hasn't published an independently-verifiable SWE-bench or Terminal-Bench 2.1 score yet.

Source: TechCrunch — OpenAI launches Astra, its powerful (and controversial) new model · CNBC — OpenAI begins rolling out Astra model after warning of its advanced cyber capabilities · OpenAI — Path to Astra: critical capabilities and frontier safeguards · Requesty — GPT-6 Astra scores 61 on the independent index, the same as Sol

Gemini 3.8 Flash — Google DeepMind (September 2, 2026)

Who / License: Google DeepMind; closed, API and Vertex AI.

What's notable: Google's fourth Flash-tier release in under four months, and still no sign of Gemini 3.5 Pro, which has now missed three announced targets since May. 3.8 Flash keeps 3.7 Flash's pricing ($0.75/$3.75 per million tokens through the end of 2026, rising to $1.50/$7.50 in January) and 1,048,576-token context, but Google says it beats 3.7 Flash on every published benchmark. Independently, Artificial Analysis has it at 81.27% on Terminal-Bench 2.1 (vals.ai) and an Intelligence Index of 58.7 — a genuinely strong score for a Flash-tier model, ahead of Qwen3.8-Max and within striking distance of GPT-5.6 Sol.

Source: Artificial Analysis — Google has released Gemini 3.8 Flash, its fourth Flash model in under four months · eesel AI — Gemini 3.8 Flash review 2026: benchmarks, pricing, and the catch

Muse Spark 1.3 — Meta (September 2, 2026)

Who / License: Meta; closed — no open weights at launch, a break from Meta's recent pattern with Muse Glimmer and the planned Muse Spark 1.2 release.

What's notable: Meta's biggest coding and agentic jump yet in the Muse line, and its highest Intelligence Index score to date — 62 (AA), good for third place this week, just behind Claude Opus 5 and ahead of GPT-6 Astra. Meta claims 75.4% on DeepSWE v1.1 and 88.8% on Terminal-Bench 2.1, both self-reported. 1M-token context, standard pricing at $1.25/$4.25 per million tokens (a "Contributor" tier at $0.10/$0.20 lets Meta train on your prompts in exchange for a steep discount). Meta says it still plans to open-weight Muse Spark 1.2; whether 1.3's weights ever follow is, per Meta's own roadmap language, undecided.

Source: VentureBeat — Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can't broadly use yet · The Register — Zuck's Muse to Spark joy with open weights release 'soon' · SiliconANGLE — Meta says it has caught up with Anthropic and OpenAI after releasing Muse Spark 1.3

Grok 4.7 slipped again: Musk's August 13 estimate of "3 to 4 weeks" has now become a September 11–12 target, alongside a claim of 2.1 trillion parameters — 40% more than Grok 4.6's 1.5 trillion. None of that is a published schedule; it's what Musk said on X on September 2.

Source: OrcaRouter — Grok 4.7 Release Date: Musk Confirms Mid-September Target

Open-weight

GLM-5.3 (full, 753B) — Zhipu / Z.ai (August 28, 2026)

Who / License: Z.ai; open weights on Hugging Face, but under a new custom "glm-5.3" license, not the plain MIT license GLM-5.2 shipped under.

What's notable: This is the flagship that missed last week's own countdown; it turned up a few days later. 753B total parameters, 40B active, 1M-token context, API pricing at $1.40/$4.40 per million tokens. GLM-5.3 reuses GLM-5.2's base — every reported gain comes from post-training alone, including a jump from 4.6 to 28.3 on Terminal-Bench 3.0. The license change is the real story: any Model-as-a-Service reseller clearing $10 billion in trailing 12-month revenue now has to pass a Z.ai security review to keep using it commercially, a restriction squarely aimed at the hyperscalers who'd otherwise resell it at scale. GLM-5.3-Flash, covered last week, still ships under plain MIT — the restriction is specific to the flagship checkpoint.

Source: The New Stack — Z.ai's GLM-5.3 goes open weight, but its new license aims at hyperscalers · Digital Applied — GLM-5.3's Weights Are Out. The License Is Not MIT · Kingy AI — GLM-5.3 Weights Are Out—But Running Them Takes Eight GPUs


Head-to-Head — Current Frontier

State of play as of September 4, 2026. AA Index = Artificial Analysis Intelligence Index, pulled from this week's leaderboard snapshot (primary). GPQA Diamond, SWE-bench, and Terminal-Bench 2.1 cells cite vals.ai's independently-run evals where marked; figures marked * are self-reported by the vendor and haven't been independently reproduced. "—" means not publicly confirmed on that specific benchmark.

Model Org Open? AA Index GPQA Diamond SWE-bench (Verified/Pro) Terminal-Bench 2.1 Context $ / Mtok in/out
Claude Fable 5.1 Anthropic No 66 (AA) 92.6% 81.2% (Pro) 85.02% (vals.ai) 1M $10 / $50
Claude Opus 5 Anthropic No 63 (AA) 84.1% 97.00% (Verified, vals.ai) 84.64% (vals.ai) 1M $5 / $25
Muse Spark 1.3 Meta No (weights undecided) 62 (AA) 88.8%* 1M $1.25 / $4.25
Claude Fable 5 Anthropic No 62 (AA) 80.0% (Pro) 88.0%* 1M $10 / $50
GPT-6 Astra OpenAI No 61 (AA) 96.0%* 1.05M $10 / $50
Grok 4.6 xAI No 60.9 (AA) 78.28% (vals.ai) 500K $2 / $6 (below 200K)
Kimi K3 Moonshot AI Yes (custom license) 59.7 (AA) 93.5% 93.40% (Verified, vals.ai) 80.90% (vals.ai) 1.05M $3 / $15
GLM-5.3 Zhipu / Z.ai Yes (custom license, MaaS clause) 59.5 (AA) 1M $1.40 / $4.40
GPT-5.6 Sol OpenAI No 58.9 (AA) 94.1% 64.6% (Pro) 85.77% (vals.ai) 1M+ $5 / $30
Gemini 3.8 Flash Google DeepMind No 58.7 (AA) 81.27% (vals.ai) 1M $0.75 / $3.75
Qwen3.8-Max (API) Alibaba No 58.1 (AA) 92.6%* 67.7%* (Pro) 86.6%* 1M $2 / $6

Reading the table: Claude Fable 5.1 didn't just win the AA Index, it won it by three points over Claude Opus 5, and it did that on day one. But "leads the index" and "leads every benchmark" aren't the same claim: Opus 5 still holds the top independently-run SWE-bench Verified score in this table (97.00%), and GPT-5.6 Sol still leads independent Terminal-Bench 2.1 (85.77%). GPT-6 Astra is the week's other headline, and its independent Intelligence Index score — 61, tied with Sol — sits noticeably below its self-reported GPQA claim of 96.0%, a gap worth watching once someone outside OpenAI runs the eval. On the open side, GLM-5.3's full weights finally landed, but its 59.5 AA Index score trails Kimi K3, and neither has closed the gap to the closed frontier this week — that gap is currently about 6.5 AA points between GLM-5.3 and Claude Fable 5.1.


Benchmark & Leaderboard Movement

  • Claude Fable 5.1 took the AA Index lead within a day of shipping, at 66 versus Claude Opus 5's 63 — the first time an Anthropic release has displaced another Anthropic release at the top of this particular leaderboard since this beat started tracking it.
  • GPT-6 Astra is the first model anywhere to cross OpenAI's "Critical" cyber capability threshold and still ship, gated behind the Daybreak program rather than delayed outright — a different resolution than the "cannot rule out critical cyber risk" holding pattern this beat reported for Astra back in early August.
  • Meta broke its own open-weight streak. After Muse Glimmer and a promised Muse Spark 1.2 release, Muse Spark 1.3 shipped closed, with Meta's own language calling any future open release "undecided."
  • GLM-5.3's full weights arrived, but the MIT era for Zhipu's flagship line ended with them. The new "glm-5.3" license adds a revenue-triggered security-review clause aimed at large-scale resellers — GLM-5.3-Flash, released two days earlier, is unaffected and stays MIT.
  • Grok 4.7 slipped again, from an August-window implication to Musk's own "3–4 weeks" (August 13) to a firmer September 11–12 target, now with a claimed 2.1T-parameter count.

Analysis

For agentic coding, this week complicates a story that had been simple for a month: Claude Fable 5.1 leads the composite index, but Claude Opus 5 still leads independent SWE-bench, and GPT-5.6 Sol still leads independent Terminal-Bench 2.1 — three different "best" answers depending which number you trust. GPT-6 Astra's self-reported coding numbers are strong, but nobody outside OpenAI has run them yet, and its independent AA score sits mid-pack. For reasoning, GPT-6 Astra's 96.0% GPQA claim is self-reported only; until someone verifies it independently, Gemini 3.1 Pro Preview's 95.45% (vals.ai, from February) is still the most trustworthy number in that column. For cheap-and-fast, Gemini 3.8 Flash is the pick this week — an AA Index of 58.7 at Flash pricing is a genuinely good trade. For open-weight and self-hosted, GLM-5.3's arrival is real, but its license is no longer the selling point GLM-5.2's was, and Kimi K3 still beats it on the composite index.

The bigger shift is structural. For most of the summer this beat led with delays — Gemini 3.5 Pro, Astra, Grok 4.7, and GLM-5.3's weights, all stuck. Two of those four unstuck this week: Astra shipped, and GLM-5.3's weights landed. Grok 4.7 at least has a firmer date now instead of a vague one. Gemini 3.5 Pro is the holdout — still no date, still three missed targets, still nothing new to report three and a half months after Google first announced it at I/O. Whether the rest is a coincidence of scheduling or three labs deciding independently to stop sitting on finished work, the standings changed more in three days than they had in the previous month.


Sources

More from News