AI Model & Benchmark Watch — September 11, 2026
After last week's four-launch pileup, the only new model this week came from DeepSeek — but Artificial Analysis quietly rewrote the scoreboard everyone else gets measured against.
AI Model & Benchmark Watch — September 11, 2026
After last week's four-launch pileup, the only new model this week came from DeepSeek — but Artificial Analysis quietly rewrote the scoreboard everyone else gets measured against.
Overview
This was a comedown week after the busiest stretch this beat has covered. DeepSeek shipped V4.1 Flash on September 10, an MIT-licensed open-weight model that claims to match Claude Opus 5 on agentic coding at roughly 1/33rd the price. Grok 4.7 missed yet another Musk deadline — it still hadn't appeared as of this writing, despite a September 11 target he set on X nine days ago. The bigger story is procedural: Artificial Analysis retired GPQA Diamond from its Intelligence Index on September 4 for being saturated and rebuilt the composite around two new evals, which dropped every model's score by 10 to 15 points overnight. None of that reflects a model getting worse — it reflects the ruler getting redrawn — but it means this week's table numbers aren't directly comparable to last week's, and we say so below rather than let a lower score imply a regression that didn't happen.
New & Updated Models (this week)
Commercial / closed
Nothing shipped this week. Grok 4.7 is still the one flagship missing from this beat's coverage — Musk's September 11–12 target, posted September 2, has come and gone without a model card, an API ID, or a price from xAI itself.
Source: Atoms — Grok 4.7 Release Date: What Elon Musk Announced and What Is Still Unconfirmed
Open-weight
DeepSeek V4.1 Flash — DeepSeek (September 10, 2026)
Who / License: DeepSeek; open weights on Hugging Face under MIT.
What's notable: A 552B-parameter mixture-of-experts model that activates just 8B parameters per token on reads (16B on generation), with a 1M-token context window and up to 384K tokens of output. DeepSeek says the redesign cuts KV cache to 890 bytes per token and adds a 1–100 reasoning-effort dial for trading accuracy against cost. Pricing is aggressive: $0.15 / $0.60 per million tokens off-peak, $0.30 / $1.20 during peak UTC hours. DeepSeek's own benchmark table — run on its own harness, so treat the decimals as a vendor's best case — has V4.1 Flash edging out Claude Opus 5 on DeepSWE v1.1 (74.2 vs. 74.0) and beating GPT-5.6 Sol on Terminal-Bench 2.1 (90.6 vs. 88.8), while trailing badly on Humanity's Last Exam (36.8 vs. Opus 5's 56.3) and on longer-horizon agentic work. Independently, Artificial Analysis puts its Intelligence Index at 40 under the new v4.2 scale — tied with Qwen3.8-Max, not with Opus 5. Multimodal input (image + text), text-only output. As of September 14, DeepSeek is also routing all deepseek-v4-pro API calls to V4.1 Flash at the Flash price, effectively retiring V4 Pro.
Source: DataNorth — DeepSeek releases DeepSeek-V4.1-Flash · Flowtivity — DeepSeek V4.1 Flash Benchmarks: Open-Weights Model Beats GPT-5.6 Sol at Agentic Coding · Artificial Analysis — DeepSeek V4.1 Flash model page
Head-to-Head — Current Frontier
State of play as of September 11, 2026. AA Index scores below use Artificial Analysis's Intelligence Index v4.2, which launched September 4 and is not comparable to the v4.1-era scores in last week's edition — the index dropped GPQA Diamond as saturated, added two private evals (AA-Briefcase, GDP.pdf), and doubled the share of held-out test sets, which pushed every score down 10–15 points regardless of model capability. GPQA Diamond and Terminal-Bench 2.1 cells cite vals.ai / Artificial Analysis's standalone runs where marked; figures marked * are self-reported by the vendor on its own harness and haven't been independently reproduced. SWE-bench Verified was archived by vals.ai on May 4, 2026 as saturated and contaminated, so its column reflects each model's last independently-run score, not a fresh weekly result — many current-generation models simply have no SWE-bench number to report. LMArena Elo is the Arena.ai (formerly LMArena) text leaderboard; "—" means the model hasn't accumulated enough votes to rank, or isn't yet listed.
| Model | Org | Open? | AA Index (v4.2) | LMArena Elo | GPQA Diamond | SWE-bench (Verified/Pro) | Terminal-Bench 2.1 | Context | $ / Mtok in/out |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | No | 53 | 1504 | 92.6% | 81.2% (Pro)* | 85.02% | 1M | $10 / $50 |
| GPT-6 Astra | OpenAI | No | 53 | — | 96.1–96.3% | — | 87.27% | 1.05M | $10 / $50 |
| Claude Opus 5 | Anthropic | No | 51 | 1493 | 84.1% | 97.00% (Verified) | 84.64% | 1M | $5 / $25 |
| Muse Spark 1.3 | Meta | No (weights undecided) | 48 | — | — | — | 79.03% | 1M | $1.25 / $4.25 |
| GPT-5.6 Sol | OpenAI | No | 47 | — | 94.1% | 64.6% (Pro)* | 85.77% | 1M+ | $5 / $30 |
| GLM-5.3 | Zhipu / Z.ai | Yes (custom license, MaaS clause) | 45 | — | — | — | — | 1M | $1.40 / $4.40 |
| Grok 4.6 | xAI | No | 44 | — | — | — | 78.28% | 500K | $2 / $6 (below 200K) |
| Kimi K3 | Moonshot AI | Yes (custom license) | 44 | 1489 | 93.5% | 93.40% (Verified) | 80.90% | 1.05M | $3 / $15 |
| Gemini 3.8 Flash | Google DeepMind | No | 41 | 1494 | 95.3% | — | 81.27% | 1M | $0.75 / $3.75 |
| Qwen3.8-Max (API) | Alibaba | No | 40 | — | 92.6%* | 67.7% (Pro)* | 86.6%* | 1M | $2 / $6 |
| DeepSeek V4.1 Flash | DeepSeek | Yes (MIT) | 40 | — | 90.9%* | — | 90.6%* | 1M | $0.15 / $0.60 |
Reading the table: Fable 5.1 and Astra are tied atop the rebuilt AA Index at 53, but they lead in different places — Astra now has the best independently-confirmed GPQA Diamond score (96.1–96.3%, close to the 96.0% OpenAI claimed at launch) and the best independent Terminal-Bench 2.1 score (87.27%, edging out GPT-5.6 Sol's long-standing 85.77%), while Fable 5.1 still leads the composite index and sits ahead of Astra on Arena's blind-preference Elo. That split is itself informative: Claude Fable 5, a model two versions old, still tops the Arena text leaderboard at 1507 — human preference and benchmark intelligence are measuring different things right now, and they don't agree on a single "best" model. On price-to-performance, DeepSeek V4.1 Flash is the story: an AA Index of 40, tied with Qwen3.8-Max, at $0.15/$0.60 per million tokens versus Qwen's $2/$6 — but that gap is still self-reported on DeepSeek's own harness for anything beyond the AA Index, so treat the coding claims as a starting point for verification, not a settled result.
Benchmark & Leaderboard Movement
- Artificial Analysis rescaled its entire Intelligence Index on September 4, retiring GPQA Diamond as saturated (24 of 136 models were already scoring 90%+) and adding AA-Briefcase and GDP.pdf, a 4,592-page long-document reasoning eval. Private, held-out test sets now make up 40% of the index weighting, double the prior version. Every model's score dropped 10–15 points as a result — Fable 5.1 went from 66 to 53, Opus 5 from 63 to 51 — with no change in the models themselves.
- GPT-6 Astra's independent benchmark numbers landed this week and mostly held up. Its GPQA Diamond score of 96.1–96.3% (AA/vals.ai) sits within a point of OpenAI's own 96.0% launch claim, and its Terminal-Bench 2.1 score of 87.27% is now the highest independently-confirmed number on that leaderboard — both were blank cells in last week's table.
- SWE-bench Verified stays frozen. Vals.ai archived it back on May 4, 2026 as saturated and contaminated (the top five models had compressed into a 93–97% band), so this week's column is the same set of stale scores as last week — Claude Opus 5's 97.00% is still the newest independently-run number on the books, five months old.
- DeepSeek V4.1 Flash is the cheapest model in this table by a wide margin — $0.15 per million input tokens is roughly a thirteenth of Qwen3.8-Max's rate and a sixty-sixth of Claude Fable 5.1's — while matching Qwen's AA Index score.
- Grok 4.7 slipped past its own deadline again. Musk's September 11–12 window, set September 2, produced no model card, no API listing, and no benchmark table as of this writing.
Analysis
For agentic coding, the honest answer this week is "it depends which number you trust." Astra now leads independent Terminal-Bench 2.1; Opus 5 still leads the frozen SWE-bench Verified board; Fable 5.1 leads the composite index. DeepSeek V4.1 Flash claims to beat all of them on a couple of agentic-loop metrics, but only on its own harness — worth trying, not worth trusting blind. For reasoning, Astra's GPQA score is now about as close to independently verified as this beat gets, which is a genuine result given OpenAI shipped it under a "Critical" cyber-capability classification and tighter access controls. For cheap-and-fast, DeepSeek V4.1 Flash is the new default to evaluate — an AA Index of 40 at $0.15/$0.60 is a real price-performance jump even before anyone outside DeepSeek confirms the coding claims. For open-weight and self-hosted, the gap to the closed frontier is still about 8 AA-Index points (GLM-5.3's 45 against Fable 5.1's 53), roughly the same gap as last week once you adjust for the rescale — open models aren't closing ground this week so much as holding it.
The more useful takeaway is a caution about the leaderboards themselves. Two of this beat's five tracked coding/reasoning benchmarks changed meaning this week without any model changing at all: GPQA Diamond got demoted from AA's composite for saturation, and SWE-bench Verified has been sitting frozen since May. When a benchmark stops differentiating models, labs and trackers replace it — which is healthy — but it also means a "score" isn't a fixed thing to compare across weeks without checking what's actually being measured. That's the whole justification for citing sources instead of quoting a number from memory.
Sources
- Atoms — Grok 4.7 Release Date: What Elon Musk Announced and What Is Still Unconfirmed
- OrcaRouter — Grok 4.7 Release Date: Musk Confirms Mid-September Target
- CellCog — Grok 4.7: Release Date, What Musk Has Promised, and What xAI Has Shipped
- DataNorth — DeepSeek releases DeepSeek-V4.1-Flash
- Flowtivity — DeepSeek V4.1 Flash Benchmarks: Open-Weights Model Beats GPT-5.6 Sol at Agentic Coding
- Artificial Analysis — DeepSeek V4.1 Flash model page
- BenchLM.ai — DeepSeek V4.1 Flash Benchmarks & Pricing (September 2026)
- Artificial Analysis — Announcing Artificial Analysis Intelligence Index v4.2
- Artificial Analysis — Artificial Analysis Intelligence Index (v4.3 methodology page)
- Artificial Analysis — LLM Leaderboard
- Artificial Analysis — GPQA Diamond Benchmark Leaderboard
- vals.ai — GPQA Diamond leaderboard
- IntuitionLabs — GPQA Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
- vals.ai — SWE-bench leaderboard
- Hacker News — I'm a co-creator of SWE-bench: SWE-bench Verified is now saturated at 93.9%
- Latent Space — The End of SWE-Bench Verified
- BenchLM.ai — Terminal-Bench 2.1 Leaderboard & Scores (September 2026)
- Arena.ai — Text leaderboard
- CNBC — 'Model fatigue' sets in as AI labs roll out new versions
- Requesty — GPT-6 Astra scores 61 on the independent index, the same as Sol
- Kingy AI — GLM-5.3 Weights Are Out—But Running Them Takes Eight GPUs
- Emergent — GLM 5.3 Benchmarks: What the Numbers Show & What They Don't
- VentureBeat — China's Moonshot AI releases Kimi K3, the largest open-source model ever
- DataCamp — Qwen3.8-Max: Features, Benchmarks, and Pricing
- My-Library — AI Model & Benchmark Watch, September 4, 2026 edition
More from News