← August 2026
News 2026-08-07

AI Model & Benchmark Watch — August 7, 2026

OpenAI says an unreleased model called Astra spent about $2,000 in compute to crack ten math problems that had sat open for a decade or more, Alibaba's Qwen3.8-Max went GA with a "second only to…

AI Model & Benchmark Watch — August 7, 2026

AI Model & Benchmark Watch — August 7, 2026

OpenAI says an unreleased model called Astra spent about $2,000 in compute to crack ten math problems that had sat open for a decade or more, Alibaba's Qwen3.8-Max went GA with a "second only to Fable 5" pitch that its own independently re-run benchmark score doesn't support, and xAI shipped Grok 4.6 with no benchmarks attached at all.

Overview

The biggest AI news of the week isn't on a leaderboard: OpenAI published Lean-verified proofs showing its next model, Astra, resolved ten longstanding open problems in math and theoretical computer science, including a 1999 question from Mikhail Gromov about non-sofic groups. Astra itself still isn't public. Closer to the ground, Alibaba pushed Qwen3.8-Max into general availability on August 3 at a genuinely aggressive $2/$6 per Mtok, but Artificial Analysis's independent score came in at 56 (after an initial run of 53 that the firm attributes to endpoint instability) — below Kimi K3's 57.1, not "second only to Fable 5" as Alibaba's own marketing claims. xAI's Grok 4.6 arrived today built on the same 1.5T parameter foundation as Grok 4.5, with the improvement coming entirely from post-training rather than scale, but xAI hasn't published a single benchmark number for it yet. Gemini 3.5 Pro, missing since a May announcement and three blown internal deadlines, is now rumored for August 12 — the same date Alibaba has set for Qwen3.8-Max's open weights.

New & Updated Models (this week)

Qwen3.8-Max — Alibaba (August 3, 2026)

Who / License: Alibaba; closed API today via Alibaba Cloud Model Studio, open weights promised August 12.

What's notable: A 2.4-trillion-parameter MoE model with 95B active parameters, 1M-token context, and native text/image/video input — Alibaba's first model over 1T parameters. Alibaba's own benchmark table claims 92.6% on GPQA Diamond and 67.7% on SWE-bench Pro (versus Claude Fable 5's 80.0%), plus a headline claim that the model completed a multi-day coding task autonomously over 16 days — a claim analyst Amit Jena pushed back on, asking how much human intervention happened along the way and whether the output survived code review. Artificial Analysis's own re-run put the Intelligence Index at 56, a 10-point jump over Qwen3.7 Max but still short of Kimi K3's 57.1 — and flagged a hallucination-rate regression on AA-Omniscience, from 23% to 40%. Alibaba also launched QwenWork, an enterprise agent platform in public beta that is subject to Chinese data-law frameworks (Cybersecurity Law, Data Security Law, PIPL) — worth a compliance review before routing regulated data through it.

Source: MarkTechPost — Alibaba Qwen Releases Qwen3.8-Max · Artificial Analysis — Qwen3.8 Max · The Decoder — Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher for 25 percent less · InfoWorld — Alibaba takes aim at OpenAI and Anthropic with Qwen3.8-Max launch · officechai — Alibaba releases Qwen 3.8 Max, beats GPT 5.6 Sol and Fable on many benchmarks · Enterprise DNA — Qwen3.8-Max is live: what it means for enterprise AI


Grok 4.6 — xAI (August 7, 2026)

Who / License: xAI; closed.

What's notable: Same 1.5T-parameter V9 foundation as Grok 4.5, with the entire capability gain coming from supervised fine-tuning and reinforcement learning rather than added scale. xAI has not published GPQA, SWE-bench, MMLU, or any other benchmark for it as of this writing — treat it as unscored until third-party evals land. A larger 2.1T Grok 4.7 is reportedly a few weeks out.

Source: kie.ai — What Is Grok 4.6? xAI's 1.5T-Param Model Explained · buildfastwithai — Grok 4.6 Preview


Astra — OpenAI (research disclosure, August 1–2, 2026; not yet released)

Who / License: OpenAI; unreleased. Sam Altman has previewed it to policymakers as a new model class alongside Sol, Terra, and Luna.

What's notable: Not a product launch — a capability disclosure. OpenAI published a 249-page manuscript and zero-"sorry" Lean 4 proof certificates on GitHub under Apache 2.0, showing Astra generated solutions to ten problems open for a decade or more across group theory, von Neumann algebras, high-dimensional geometry, quantum complexity, lattice cryptography, and extremal combinatorics — including an explicit construction of a non-sofic group, open since Gromov posed the question in 1999, and a disproof of Connes's rigidity conjecture. OpenAI says the compute cost was about $2,000. Thomas Bloom, the mathematician who debunked a false AI math claim in 2025, called it "big news" — distinguishing it from earlier hype cycles because the proofs are machine-checked, not just asserted.

Source: SiliconANGLE — OpenAI's Astra solves 10 long-open math problems and publishes the proofs · Tech Times — OpenAI's Astra Solves Ten Decade-Old Math Problems With Machine-Checkable Lean Proofs · Forbes — OpenAI's Astra Solved Decades-Old Math Problems For $2,000 · DataCamp — OpenAI's New Model, Astra, Has Solved Ten Open Math Problems


Muse Spark 1.2 & Muse Code — Meta (August 5, 2026)

Who / License: Meta Superintelligence Labs; closed, API in public preview.

What's notable: A coding-focused update to July's Muse Spark 1.1, shipped alongside Muse Code, a beta terminal coding agent co-trained with the model. 1M-token context, built for large-repo work and long-running multi-agent development sessions. Standard pricing is $1.25 input / $4.25 output per Mtok; a new "contributor" tier drops to $0.10 / $0.20 in exchange for letting Meta train on your prompts and completions, capped at 60 requests per minute.

Source: MarkTechPost — Meta AI Releases Muse Code (Beta) · Unite.AI — Meta Ships Muse Code Coding Agent With Co-Trained Muse Spark 1.2 Model · Yahoo Finance — Meta debuts Muse Spark 1.2 and first coding agent


DeepSeek V4-Flash-0731 — DeepSeek (July 31, 2026)

Who / License: DeepSeek; open weights.

What's notable: A re-post-trained refresh of V4-Flash aimed at coding and agent workflows — 284B total parameters, 13B active. DeepSeek reports 82.7 on Terminal-Bench 2.1 and 70.3 on Toolathlon Verified. Pricing held at $0.14/M cache-miss input and $0.28/M output, among the cheapest frontier-adjacent options available.

Source: MarkTechPost — DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains · Artificial Analysis — DeepSeek V4 Flash 0731


Head-to-Head — Current Frontier

The table reflects the state of play as of August 7, 2026. AA Index = Artificial Analysis Intelligence Index v4.1.1 (independently measured); scores marked * are self-reported by the vendor, not yet an Artificial Analysis entry. SWE-bench Pro throughout — not comparable to Verified scores; footnoted where only Verified is published. "—" = not publicly confirmed or not yet evaluated.

Model Org Open? AA Index GPQA Dia SWE-bench Pro Terminal-Bench 2.1 Context $ / M in/out
Claude Opus 5 Anthropic No 60.7 (#1) 93.7% 79.2% 89.1% 1M $5 / $25
Claude Fable 5 Anthropic No 59.9 (#2) 80.0% 88.0% 1M $10 / $50
GPT-5.6 Sol OpenAI No 58.9 (#3) 94.1% 64.6% 88.8% / 91.9%† 1M+ $5 / $30
Kimi K3 Moonshot AI Yes‡ 57.1 (#4) 93.5% — (76.8% Verified) 88.3% 1M $3 / $15
Qwen3.8-Max Alibaba No§ 56.0 (#5) 92.6%* 67.7%* 86.6%* 1M $2 / $6
Claude Opus 4.8 Anthropic No 55.7 (#6) 69.2% ~85.0% 1M $5 / $25
GPT-5.6 Terra OpenAI No 55.0 (#7) 63.4% 87.4% 1M+ $2 / $12
Grok 4.5 xAI No 53.8 (#9) 93.0% 64.7% 500K $2 / $6
GLM-5.2 Zhipu / Z.ai Yes (MIT) 51.1 91.2% 62.1% 82.7% 1M $1.40 / $4.40
DeepSeek V4-Pro DeepSeek Yes (MIT) 44.0 — (80.6% Verified) 1M $0.44 / $0.87

*Qwen3.8-Max: GPQA Diamond, SWE-bench Pro, and Terminal-Bench figures are Alibaba's own reported numbers, not independently verified — treat with the same caution Alibaba's "second only to Fable 5" claim deserves. Its Artificial Analysis Intelligence Index score (56.0) is independently measured. †GPT-5.6 Sol (standard): 88.8%; Sol Ultra (4-agent Codex mode): 91.9%. ‡Kimi K3: full open weights shipped July 26 under a custom "Kimi K3 License" (not MIT), with a commercial-scale display requirement. §Qwen3.8-Max: closed API today; open weights due August 12, license not yet named.

Reading the table: Grok 4.6 launched today but isn't in this table — xAI hasn't published a single benchmark for it, so there's nothing to compare yet. Qwen3.8-Max slots in at #5 on the independently-measured AA Index, ahead of Claude Opus 4.8 but behind Kimi K3, which undercuts Alibaba's own marketing pitch. It's also the cheapest closed frontier-tier model in the table at $2/$6 — cheaper than Grok 4.5 on output and a third of GPT-5.6 Sol's rate — which is probably the more durable story than its contested benchmark claims. Anthropic still holds the top two AA Index spots and the Terminal-Bench and SWE-bench Pro leads. GPQA Diamond remains saturated in the low-to-mid 90s across six different labs (Sol 94.1%, Opus 5 93.7%, Kimi K3 93.5%, Grok 4.5 93.0%, Qwen3.8-Max 92.6% self-reported, GLM-5.2 91.2%) — it stopped differentiating frontier models months ago.


Benchmark & Leaderboard Movement

  • Qwen3.8-Max's Artificial Analysis score moved twice in one week. The first run scored 53, which Artificial Analysis attributed to instability on the tested endpoint; a re-run on Alibaba's public API on August 6 brought it to 56 — still short of Kimi K3, contradicting Alibaba's "second only to Fable 5" framing from its July 19 preview.
  • Qwen3.8-Max's hallucination rate nearly doubled on Artificial Analysis's AA-Omniscience knowledge check, from 23% to 40%, alongside the intelligence-score bump — a reminder that a rising index score doesn't move every underlying metric in the same direction.
  • SWE-bench Verified is now crowded at the very top of the Anthropic lineup: Claude Opus 5 leads at 96%, with Claude Mythos 5 (95.5%, the restricted cybersecurity/biosecurity variant sharing Fable 5's base weights) and Claude Fable 5 (95%) close behind — the top three spots span less than a point.
  • Kimi K3 landed in GitHub Copilot on August 6, expanding distribution for Moonshot's open-weight model (shipped July 26) at $3/$15 per Mtok with a $0.30 cached-input rate.
  • Gemini 3.5 Pro is still missing after three blown internal deadlines (June, July, and a widely reported July 17 date); it's now rumored for August 12, the same date set for Qwen3.8-Max's open weights and reportedly close to Zhipu's still-unconfirmed GLM-5.5.
  • GLM-5.5 remains unconfirmed. Community leaks describe a 1T+-parameter model with open weights targeting an August release, but Zhipu's official channels still list GLM-5.2 as current, with no model card or benchmark published.

Analysis

For agentic coding, Claude Opus 5 and Fable 5 still set the pace, and nothing this week changed that — Qwen3.8-Max's self-reported SWE-bench Pro score (67.7%) trails Fable 5's verified 80.0% by a wide margin, whatever the autonomous-coding marketing claims suggest. For reasoning work, six labs are now within three points of each other on GPQA Diamond in the low-to-mid 90s; pick on price and context rather than chasing a benchmark that's stopped separating the field. For cheap-and-fast, DeepSeek V4-Flash-0731 ($0.14/$0.28) and Qwen3.8-Max ($2/$6) both got more competitive this week, though only the DeepSeek number comes with a verifiable track record — Qwen3.8-Max's headline claims are still Alibaba's own. For open-weight self-hosting, Kimi K3 remains the strongest fully verified option; Qwen3.8-Max's own weights aren't out until August 12, and its license terms aren't named yet.

The open-vs-closed gap didn't move much this week, but the credibility gap did: Qwen3.8-Max is the second Chinese lab release in three weeks (after Kimi K3) to launch with self-reported numbers that outpace what independent measurement actually finds, and QwenWork's Chinese-law data exposure is a real diligence item for regulated enterprises evaluating it. Meanwhile the most consequential capability news of the week — Astra's math proofs — came from a model nobody outside OpenAI can test yet, which is its own kind of unverifiable.


Sources

More from News