← August 2026
News 2026-08-28

AI Model & Benchmark Watch — August 28, 2026

Today is the day Z.ai promised GLM-5.3's full open weights, and as of publication they still hadn't shown up on Hugging Face — but two smaller open models from Z.ai and Alibaba shipped anyway, both…

AI Model & Benchmark Watch — August 28, 2026

AI Model & Benchmark Watch — August 28, 2026

Today is the day Z.ai promised GLM-5.3's full open weights, and as of publication they still hadn't shown up on Hugging Face — but two smaller open models from Z.ai and Alibaba shipped anyway, both posting numbers in Claude Opus territory at a fraction of the size.

Overview

The three delayed flagships from last week (Gemini 3.5 Pro, Astra, Grok 4.7) are still delayed, with nothing new to report beyond the date slipping further out of reach. What actually moved this week happened one tier down: Z.ai released GLM-5.3-Flash (the model that had been running anonymously as "Ox Alpha"), Alibaba released Qwen3.8-Flash-Next as an early look at its next architecture, and both landed benchmark claims that beat Claude Opus 4.8 on agentic coding tasks. Neither has independent verification yet, and neither model is the one everyone's actually waiting on: the full 744B GLM-5.3, whose open-weight release was due around today and hasn't appeared.

New & Updated Models (this week)

Commercial / closed

Anthropic shipped no new Claude model this week. The "Fable 5.1" rumor covered in recent editions is still just that: a rumor. As of August 27, Anthropic's model catalog, pricing page, and release notes list Claude Fable 5, not a 5.1. The closest thing to evidence is a "Model 2" mentioned in Anthropic's own risk report, which some outlets speculate could be an internal Fable 5.1 build, but Anthropic hasn't confirmed the connection. Source: AIToolsReview — Claude Fable 5.1 Fact Check: What's Actually Confirmed

Open-weight

GLM-5.3-Flash ("Ox Alpha") — Zhipu / Z.ai (August 26, 2026)

Who / License: Z.ai; open weights, MIT license, on Hugging Face.

What's notable: This is the model that had been running under the stealth handle "Ox Alpha" for weeks; Z.ai confirmed the identity and shipped weights the same day. It's a 320B-parameter mixture-of-experts model with only 18B active per token, natively multimodal, with a 1M-token context window. Z.ai's own numbers put it at 84.3 on Terminal-Bench 2.1, close to Claude Opus 4.8's 85.0, and claim it beats Opus 4.8 on DeepSWE (63.4) and AutomationBench (48.8) while costing roughly a tenth as much to serve. All of that is vendor-reported on Z.ai's own harness; nothing here is independently verified yet. It is not the full GLM-5.3 (see below); it's a separate, smaller checkpoint.

Source: MarkTechPost — Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE · explainx.ai — GLM-5.3-Flash Launch — Ox Alpha Was Zhipu (MIT)

Qwen3.8-Flash-Next — Alibaba (August 24–26, 2026)

Who / License: Alibaba's Tongyi Lab; open weights on Hugging Face (August 24) and ModelScope (August 26).

What's notable: Billed as an early preview of Alibaba's next architecture generation, this is a 125B-parameter MoE with only 6B active per token, plus a 51B n-gram embedding table and a 4B multi-token-prediction layer, an unusual design even by MoE standards. It reads images natively and carries a 262K-token context window. Alibaba's own benchmark card claims it beats Claude Opus 4.6 Max on SWE-bench Pro (62.5 vs. 53.4), CoWorkBench (73.9 vs. 68.2), and JobBench (55.7 vs. 36.6), plus 91.7 on GPQA Diamond and 91.9 on LiveCodeBench v6. Again: all vendor-reported, on Qwen's own harness, and not yet reproduced by a third party.

Source: explainx.ai — Qwen3.8-Flash-Next 125B-A6B: MoE Drop Aug 26 · DataCamp — Qwen3.8-Flash-Next: Features, Benchmarks, and Pricing

DeepSeek V4 Flash Vision Exp — DeepSeek (August 21, 2026)

Who / License: DeepSeek; API access, MIT-lineage base model.

What's notable: An experimental vision-enabled build of DeepSeek V4 Flash 0731 — adds image input (screenshots, charts, documents) while matching the text-only model on reasoning and agent tasks. 13B active parameters out of 284B total, 1M-token context, and it bills at the same V4 Flash rate: $0.14 per million input tokens, $0.28 per million output, images tokenized at up to 384 tokens each.

Source: DeepSeek API Docs — DeepSeek-V4-Flash-Vision-Exp Release · Digital Applied — DeepSeek V4 Flash Vision: Images, Same Price, Two Clocks

GLM-5.3 (full, 744B) open weights — still not out

Who / License: Z.ai; API has been live since August 14, open weights promised roughly two weeks later.

What's notable: Z.ai's own Hugging Face placeholder counted down to today, August 28, for the full GLM-5.3's open weights, tied to a safety review after post-training reportedly gave the model exploit-chain reasoning nobody had intentionally trained in (Z.ai says it found 1,097 critical vulnerabilities in Linux, WebKit, and FreeBSD during testing). As of this writing, the weights have not appeared. GLM-5.3-Flash, covered above, is not a substitute; it's a different, smaller model on a separately trained base.

Source: Kingy AI — GLM-5.3's Open-Weight Reality Check · MindStudio — When Will GLM 5.3 Open Weights Be Released?


Head-to-Head — Current Frontier

State of play as of August 28, 2026. AA Index = Artificial Analysis Intelligence Index, pulled this week directly from Artificial Analysis's leaderboard (primary). GPQA Diamond and SWE-bench/Terminal-Bench cells cite vals.ai's independently-run evals where marked, or are self-reported by the vendor where marked *. "—" means not publicly confirmed or not evaluated on that benchmark. GLM-5.3-Flash and Qwen3.8-Flash-Next, both new this week, aren't in the table below — neither has an independent AA Index score yet, and both sets of headline numbers are vendor-reported only; see the New Models section above.

Model Org Open? AA Index GPQA Diamond SWE-bench (Verified/Pro) Terminal-Bench 2.1 Context $ / Mtok in/out
Claude Opus 5 Anthropic No 63 (AA, primary) 84.1% 97.00% (Verified, vals.ai) 84.64% (vals.ai) 1M $5 / $25
Claude Fable 5 Anthropic No 62 (AA, primary) 80.0% (Pro) 88.0% 1M $10 / $50
Grok 4.6 xAI No 61 (AA, primary) 88.4% (AA) 500K $2 / $6 (below 200K)
GPT-5.6 Sol OpenAI No 61 (AA, primary) 94.1% 64.6% (Pro) 85.77% (vals.ai) 1M+ $5 / $30
Kimi K3 Moonshot AI Yes (Modified MIT) 60 (AA, primary) 93.5% 93.40% (Verified, vals.ai) 80.90% (vals.ai) 1.05M $3 / $15
GLM-5.3 Zhipu / Z.ai Weights promised today, not yet out 60 (AA, primary) — (91.2% was GLM-5.2's) 1M $1.40 / $4.40
Claude Opus 4.8 Anthropic No 55.7 69.2% (Pro) ~85.0% 1M $5 / $25
Qwen3.8-Max (API) Alibaba No (open weights are a separate, stripped-down checkpoint) 58 (AA, primary) 92.6%* 67.7%* (Pro) 86.6%* 1M $2 / $6
DeepSeek V4-Pro-0813 DeepSeek MIT lineage; not yet on Hugging Face 53 90.1%* 96.40% (Verified, vals.ai) vs. 80.6%* self-reported 1M $0.44 / $0.87
Gemini 3.1 Pro Preview Google DeepMind No 95.45% (vals.ai, GPQA leader)

Reading the table: Claude Opus 5 hasn't been challenged in three straight weeks of this beat. It still holds the top AA Index score and both independently-run coding evals in this table, and nothing this week changed that; the two new open models (GLM-5.3-Flash, Qwen3.8-Flash-Next) are targeting a different comparison, Opus 4.8 rather than Opus 5, and doing it on vendor-reported numbers that haven't been checked by anyone outside the lab that made them. GPQA Diamond is still led, oddly, by Gemini 3.1 Pro Preview, a model from February. Google hasn't shipped anything since to replace it. DeepSeek's SWE-bench gap (96.40% independent vs. 80.6% self-reported) remains unresolved.


Benchmark & Leaderboard Movement

  • GLM-5.3's promised open-weight date arrived without the weights. Z.ai's own Hugging Face countdown pointed to today; as of publication, nothing's there. The safety review tied to the model's unplanned exploit-chain capability is the stated reason, and there's no new date to replace the missed one.
  • Two open efficiency plays launched instead, both claiming to beat Claude Opus 4.8 on agentic coding — GLM-5.3-Flash (320B/18B active, MIT) and Qwen3.8-Flash-Next (125B/6B active, previewing Alibaba's next architecture). Read both claims as marketing until a third party reruns them; the pattern of "vendor-reported number beats a specific competitor" is exactly what this beat treats with the most caution.
  • Grok 4.7's window kept sliding. Musk's "3–4 weeks" from August 13 now lands in early September rather than the August target xAI had implied earlier in the month.
  • Fable 5.1 is still unconfirmed. No model card, no pricing page, no benchmark entry — just a "Model 2" reference in Anthropic's own risk report that outlets are guessing about.
  • Google quietly shipped a real product, Gemini 3.5 Transcribe, while its flagship LLM remains stuck in preview. Worth noting as a reminder that "no Pro model" doesn't mean "no releases."

Analysis

For agentic coding, Claude Opus 5 keeps leading on every independently-run eval in this table, and that hasn't moved in three weeks — the loudest claims against it this week (from GLM-5.3-Flash and Qwen3.8-Flash-Next) are both self-reported and both compare against Opus 4.8, not Opus 5. For reasoning, GPQA Diamond's leaderboard is still frozen on a six-month-old Gemini preview, which says more about Google's release cadence than about who's actually ahead. For cheap-and-fast, GLM-5.3's API pricing at $1.40/$4.40 remains one of the best value plays at its AA Index tier, though its self-hostable weights are still nowhere to be found. For local and self-hosted, this week's two open releases are worth watching but not yet worth trusting — wait for someone outside Z.ai or Alibaba to run the same evals before treating either as a real Opus-4.8-class contender.

The broader shape hasn't changed: three closed labs are each sitting on an unreleased flagship, and the open-weight side keeps chipping away with smaller, cheaper releases rather than head-on wins at the top. This week added a new wrinkle: an open lab missing its own self-imposed deadline, a smaller and quieter delay than what's happening at Google, OpenAI, and xAI, but the same basic story. Readiness keeps losing to ambition.


Sources

More from News