AI Model & Benchmark Watch — August 28, 2026
Today is the day Z.ai promised GLM-5.3's full open weights, and as of publication they still hadn't shown up on Hugging Face — but two smaller open models from Z.ai and Alibaba shipped anyway, both…
AI Model & Benchmark Watch — August 28, 2026
Today is the day Z.ai promised GLM-5.3's full open weights, and as of publication they still hadn't shown up on Hugging Face — but two smaller open models from Z.ai and Alibaba shipped anyway, both posting numbers in Claude Opus territory at a fraction of the size.
Overview
The three delayed flagships from last week (Gemini 3.5 Pro, Astra, Grok 4.7) are still delayed, with nothing new to report beyond the date slipping further out of reach. What actually moved this week happened one tier down: Z.ai released GLM-5.3-Flash (the model that had been running anonymously as "Ox Alpha"), Alibaba released Qwen3.8-Flash-Next as an early look at its next architecture, and both landed benchmark claims that beat Claude Opus 4.8 on agentic coding tasks. Neither has independent verification yet, and neither model is the one everyone's actually waiting on: the full 744B GLM-5.3, whose open-weight release was due around today and hasn't appeared.
New & Updated Models (this week)
Commercial / closed
- Gemini 3.5 Transcribe — Google, closed, announced August 26. A speech-to-text model, not a general-purpose LLM, so it doesn't belong in the head-to-head table below, but it's a real ship: public preview in the Gemini API, 85+ language detection, and a reported 2.6% word error rate non-streaming / 4.0% streaming, measured by Artificial Analysis. It already powers the Rambler dictation feature on Pixel 11 phones and is headed to Chrome. Source: Google — Intelligent transcription with Gemini 3.5 Transcribe · MarkTechPost — Google AI Releases Gemini 3.5 Transcribe
- Gemini 3.5 Pro remains stuck in limited Vertex AI enterprise preview, now past three missed targets (late June, July 17, early August) with no fourth date offered. Source: The AI Rankings — Gemini 3.5 Pro Release Date: Three Delays and Still Unreleased
- Astra has had no update since OpenAI's August 7 disclosure that it "cannot rule out critical cyber risk" under its Preparedness Framework. Still no release date. Source: Axios — OpenAI slows release of Astra model citing cyber capabilities
- Grok 4.7 slipped again: Musk's own "3 to 4 weeks" estimate, given August 13, now points to early September instead of the August window xAI had been signaling. Source: OrcaRouter — Grok 4.7 Release Date: 3–4 Weeks, and Already Slipping
Anthropic shipped no new Claude model this week. The "Fable 5.1" rumor covered in recent editions is still just that: a rumor. As of August 27, Anthropic's model catalog, pricing page, and release notes list Claude Fable 5, not a 5.1. The closest thing to evidence is a "Model 2" mentioned in Anthropic's own risk report, which some outlets speculate could be an internal Fable 5.1 build, but Anthropic hasn't confirmed the connection. Source: AIToolsReview — Claude Fable 5.1 Fact Check: What's Actually Confirmed
Open-weight
GLM-5.3-Flash ("Ox Alpha") — Zhipu / Z.ai (August 26, 2026)
Who / License: Z.ai; open weights, MIT license, on Hugging Face.
What's notable: This is the model that had been running under the stealth handle "Ox Alpha" for weeks; Z.ai confirmed the identity and shipped weights the same day. It's a 320B-parameter mixture-of-experts model with only 18B active per token, natively multimodal, with a 1M-token context window. Z.ai's own numbers put it at 84.3 on Terminal-Bench 2.1, close to Claude Opus 4.8's 85.0, and claim it beats Opus 4.8 on DeepSWE (63.4) and AutomationBench (48.8) while costing roughly a tenth as much to serve. All of that is vendor-reported on Z.ai's own harness; nothing here is independently verified yet. It is not the full GLM-5.3 (see below); it's a separate, smaller checkpoint.
Source: MarkTechPost — Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE · explainx.ai — GLM-5.3-Flash Launch — Ox Alpha Was Zhipu (MIT)
Qwen3.8-Flash-Next — Alibaba (August 24–26, 2026)
Who / License: Alibaba's Tongyi Lab; open weights on Hugging Face (August 24) and ModelScope (August 26).
What's notable: Billed as an early preview of Alibaba's next architecture generation, this is a 125B-parameter MoE with only 6B active per token, plus a 51B n-gram embedding table and a 4B multi-token-prediction layer, an unusual design even by MoE standards. It reads images natively and carries a 262K-token context window. Alibaba's own benchmark card claims it beats Claude Opus 4.6 Max on SWE-bench Pro (62.5 vs. 53.4), CoWorkBench (73.9 vs. 68.2), and JobBench (55.7 vs. 36.6), plus 91.7 on GPQA Diamond and 91.9 on LiveCodeBench v6. Again: all vendor-reported, on Qwen's own harness, and not yet reproduced by a third party.
Source: explainx.ai — Qwen3.8-Flash-Next 125B-A6B: MoE Drop Aug 26 · DataCamp — Qwen3.8-Flash-Next: Features, Benchmarks, and Pricing
DeepSeek V4 Flash Vision Exp — DeepSeek (August 21, 2026)
Who / License: DeepSeek; API access, MIT-lineage base model.
What's notable: An experimental vision-enabled build of DeepSeek V4 Flash 0731 — adds image input (screenshots, charts, documents) while matching the text-only model on reasoning and agent tasks. 13B active parameters out of 284B total, 1M-token context, and it bills at the same V4 Flash rate: $0.14 per million input tokens, $0.28 per million output, images tokenized at up to 384 tokens each.
Source: DeepSeek API Docs — DeepSeek-V4-Flash-Vision-Exp Release · Digital Applied — DeepSeek V4 Flash Vision: Images, Same Price, Two Clocks
GLM-5.3 (full, 744B) open weights — still not out
Who / License: Z.ai; API has been live since August 14, open weights promised roughly two weeks later.
What's notable: Z.ai's own Hugging Face placeholder counted down to today, August 28, for the full GLM-5.3's open weights, tied to a safety review after post-training reportedly gave the model exploit-chain reasoning nobody had intentionally trained in (Z.ai says it found 1,097 critical vulnerabilities in Linux, WebKit, and FreeBSD during testing). As of this writing, the weights have not appeared. GLM-5.3-Flash, covered above, is not a substitute; it's a different, smaller model on a separately trained base.
Source: Kingy AI — GLM-5.3's Open-Weight Reality Check · MindStudio — When Will GLM 5.3 Open Weights Be Released?
Head-to-Head — Current Frontier
State of play as of August 28, 2026. AA Index = Artificial Analysis Intelligence Index, pulled this week directly from Artificial Analysis's leaderboard (primary). GPQA Diamond and SWE-bench/Terminal-Bench cells cite vals.ai's independently-run evals where marked, or are self-reported by the vendor where marked *. "—" means not publicly confirmed or not evaluated on that benchmark. GLM-5.3-Flash and Qwen3.8-Flash-Next, both new this week, aren't in the table below — neither has an independent AA Index score yet, and both sets of headline numbers are vendor-reported only; see the New Models section above.
| Model | Org | Open? | AA Index | GPQA Diamond | SWE-bench (Verified/Pro) | Terminal-Bench 2.1 | Context | $ / Mtok in/out |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | No | 63 (AA, primary) | 84.1% | 97.00% (Verified, vals.ai) | 84.64% (vals.ai) | 1M | $5 / $25 |
| Claude Fable 5 | Anthropic | No | 62 (AA, primary) | — | 80.0% (Pro) | 88.0% | 1M | $10 / $50 |
| Grok 4.6 | xAI | No | 61 (AA, primary) | — | — | 88.4% (AA) | 500K | $2 / $6 (below 200K) |
| GPT-5.6 Sol | OpenAI | No | 61 (AA, primary) | 94.1% | 64.6% (Pro) | 85.77% (vals.ai) | 1M+ | $5 / $30 |
| Kimi K3 | Moonshot AI | Yes (Modified MIT) | 60 (AA, primary) | 93.5% | 93.40% (Verified, vals.ai) | 80.90% (vals.ai) | 1.05M | $3 / $15 |
| GLM-5.3 | Zhipu / Z.ai | Weights promised today, not yet out | 60 (AA, primary) | — (91.2% was GLM-5.2's) | — | — | 1M | $1.40 / $4.40 |
| Claude Opus 4.8 | Anthropic | No | 55.7 | — | 69.2% (Pro) | ~85.0% | 1M | $5 / $25 |
| Qwen3.8-Max (API) | Alibaba | No (open weights are a separate, stripped-down checkpoint) | 58 (AA, primary) | 92.6%* | 67.7%* (Pro) | 86.6%* | 1M | $2 / $6 |
| DeepSeek V4-Pro-0813 | DeepSeek | MIT lineage; not yet on Hugging Face | 53 | 90.1%* | 96.40% (Verified, vals.ai) vs. 80.6%* self-reported | — | 1M | $0.44 / $0.87 |
| Gemini 3.1 Pro Preview | Google DeepMind | No | — | 95.45% (vals.ai, GPQA leader) | — | — | — | — |
Reading the table: Claude Opus 5 hasn't been challenged in three straight weeks of this beat. It still holds the top AA Index score and both independently-run coding evals in this table, and nothing this week changed that; the two new open models (GLM-5.3-Flash, Qwen3.8-Flash-Next) are targeting a different comparison, Opus 4.8 rather than Opus 5, and doing it on vendor-reported numbers that haven't been checked by anyone outside the lab that made them. GPQA Diamond is still led, oddly, by Gemini 3.1 Pro Preview, a model from February. Google hasn't shipped anything since to replace it. DeepSeek's SWE-bench gap (96.40% independent vs. 80.6% self-reported) remains unresolved.
Benchmark & Leaderboard Movement
- GLM-5.3's promised open-weight date arrived without the weights. Z.ai's own Hugging Face countdown pointed to today; as of publication, nothing's there. The safety review tied to the model's unplanned exploit-chain capability is the stated reason, and there's no new date to replace the missed one.
- Two open efficiency plays launched instead, both claiming to beat Claude Opus 4.8 on agentic coding — GLM-5.3-Flash (320B/18B active, MIT) and Qwen3.8-Flash-Next (125B/6B active, previewing Alibaba's next architecture). Read both claims as marketing until a third party reruns them; the pattern of "vendor-reported number beats a specific competitor" is exactly what this beat treats with the most caution.
- Grok 4.7's window kept sliding. Musk's "3–4 weeks" from August 13 now lands in early September rather than the August target xAI had implied earlier in the month.
- Fable 5.1 is still unconfirmed. No model card, no pricing page, no benchmark entry — just a "Model 2" reference in Anthropic's own risk report that outlets are guessing about.
- Google quietly shipped a real product, Gemini 3.5 Transcribe, while its flagship LLM remains stuck in preview. Worth noting as a reminder that "no Pro model" doesn't mean "no releases."
Analysis
For agentic coding, Claude Opus 5 keeps leading on every independently-run eval in this table, and that hasn't moved in three weeks — the loudest claims against it this week (from GLM-5.3-Flash and Qwen3.8-Flash-Next) are both self-reported and both compare against Opus 4.8, not Opus 5. For reasoning, GPQA Diamond's leaderboard is still frozen on a six-month-old Gemini preview, which says more about Google's release cadence than about who's actually ahead. For cheap-and-fast, GLM-5.3's API pricing at $1.40/$4.40 remains one of the best value plays at its AA Index tier, though its self-hostable weights are still nowhere to be found. For local and self-hosted, this week's two open releases are worth watching but not yet worth trusting — wait for someone outside Z.ai or Alibaba to run the same evals before treating either as a real Opus-4.8-class contender.
The broader shape hasn't changed: three closed labs are each sitting on an unreleased flagship, and the open-weight side keeps chipping away with smaller, cheaper releases rather than head-on wins at the top. This week added a new wrinkle: an open lab missing its own self-imposed deadline, a smaller and quieter delay than what's happening at Google, OpenAI, and xAI, but the same basic story. Readiness keeps losing to ambition.
Sources
- The AI Rankings — Gemini 3.5 Pro Release Date: Three Delays and Still Unreleased
- codersera — Gemini 3.5 Pro Release Date: Why It's Delayed
- Axios — Exclusive: OpenAI slows release of Astra model citing cyber capabilities
- OpenAI — Responding to the next frontier of critical cyber capabilities
- OrcaRouter — Grok 4.7 Release Date: 3–4 Weeks, and Already Slipping
- techjournal.org — Grok 4.7 Release Date: Why xAI Delayed It to September
- AIToolsReview — Claude Fable 5.1 Fact Check: What's Actually Confirmed (August 2026)
- evolink.ai — Claude Fable 5.1 Release Date, Status & What We Know
- Anthropic — Claude Fable 5 and Claude Mythos 5
- Google — Intelligent transcription with Gemini 3.5 Transcribe
- MarkTechPost — Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages
- 9to5Google — Google launches Gemini 3.5 Transcribe, which powers Gboard Rambler & is coming to Chrome
- MarkTechPost — Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
- explainx.ai — GLM-5.3-Flash Launch — Ox Alpha Was Zhipu (MIT)
- Kingy AI — Ox Alpha Confirmed as GLM-5.3 Flash: Specs & Price
- Kingy AI — GLM-5.3's Open-Weight Reality Check: The Two-Week Delay, Cybersecurity Leap, and 2,436-Vulnerability Claim
- MindStudio — When Will GLM 5.3 Open Weights Be Released?
- Distk — GLM-5.3 in 2026: The Open Model That Shipped Without Its Weights
- techtimes — GLM-5.3: Post-Training Produced Exploit Chains Z.ai Never Planned, Finds 1,097 Critical Bugs
- explainx.ai — Qwen3.8-Flash-Next 125B-A6B: MoE Drop Aug 26 (2026)
- DataCamp — Qwen3.8-Flash-Next: Features, Benchmarks, and Pricing
- OrcaRouter — Qwen3.8-Flash-Next Is Out: Qwen4 Architecture Confirmed
- Bloomberg — Alibaba Releases Smaller Qwen AI Model to Compete With Anthropic, DeepSeek
- DeepSeek API Docs — DeepSeek-V4-Flash-Vision-Exp Release: Multimodal API Now Live
- Digital Applied — DeepSeek V4 Flash Vision: Images, Same Price, Two Clocks
- eesel AI — Claude Opus 5 pricing in 2026: API costs, plans, real bills
- ayautomate — Claude Opus 5 Pricing: $5 and $25 per Million Tokens
- Artificial Analysis — LLM Leaderboard
- vals.ai — SWE-bench leaderboard
- vals.ai — GPQA leaderboard
- vals.ai — Terminal-Bench 2.1 leaderboard
- localaimaster — LMArena Leaderboard (Live): Top AI Models Ranked
- swfte — LMArena.ai Top Models August 2026 | LMSys Rebrand & Current Leaders
- emergent.sh — GLM 5.3 Benchmarks: What the Numbers Show & What They Don't
- morphllm — Kimi K3: 2.8T Parameters, Open Weights, 1M Context, Benchmarks, Pricing
- AIToolsReview — Kimi K4 Release Date: What's Confirmed vs Rumoured (August 2026)
More from News