AI Model & Benchmark Watch — July 10, 2026
GPT-5.6 Sol, Terra, and Luna went fully public on July 9, instantly seizing the #2 slot on the Artificial Analysis Intelligence Index and setting a new Terminal-Bench record — while OpenAI…
AI Model & Benchmark Watch — July 10, 2026
GPT-5.6 Sol, Terra, and Luna went fully public on July 9, instantly seizing the #2 slot on the Artificial Analysis Intelligence Index and setting a new Terminal-Bench record — while OpenAI simultaneously shipped full-duplex voice models and a workplace agent product in its most coordinated product blitz of the year.
Overview
The two days since the July 8 edition have seen the most concentrated OpenAI product release in years. GPT-5.6 (all three tiers) cleared its U.S. government-gated partner preview and went globally available on July 9 across ChatGPT, Codex, and the OpenAI API. On the same day, OpenAI launched ChatGPT Work — a long-horizon agentic workplace product powered by GPT-5.6 and Codex. One day earlier (July 8), OpenAI introduced GPT-Live-1 and GPT-Live-1 mini, full-duplex voice models that listen and speak simultaneously. GPT-5.6 Sol now sits at #2 on the Artificial Analysis Intelligence Index (score 59, vs. Claude Fable 5's 60) and sets a new Terminal-Bench 2.1 state-of-the-art at 91.9% in Sol Ultra mode. Claude Fable 5 holds the #1 index position and still leads on SWE-bench Pro (80.3% vs. Sol's 64.6%), keeping it as the definitive agentic coding benchmark leader — for now.
On the open-weight front, no new releases this week. GLM-5.2 (MIT, July debut) and DeepSeek V4-Pro remain the headline open models. Gemini 3.5 Pro continues to slip, with reports now pointing to a July 17 target that Google has not officially confirmed.
New & Updated Models (this week — July 8–10)
GPT-5.6 Sol / Terra / Luna — OpenAI (GA: July 9, 2026)
Who / License: OpenAI; closed, proprietary.
What's notable: Three-tier family that previewed June 26 and spent two weeks in a U.S. government-gated partner rollout before going public on July 9 globally across ChatGPT, Codex, and the API. Sol is the frontier flagship; Terra delivers Sol-adjacent performance (~87% on Terminal-Bench) at 2× lower cost; Luna is the fastest and cheapest. Sol Fast, an ultra-premium serving option on Cerebras hardware ($12.50/$75 per M tokens), delivers up to 750 tokens/second for latency-critical pipelines. Context: up to 1M+ tokens (preview materials cited 1.5M; GA API docs vary by tier). The headline benchmark claim is Sol Ultra's 91.9% on Terminal-Bench 2.1 (4-agent Codex mode) — the highest published figure on that benchmark, beating Claude Fable 5's 86.0%. On SWE-bench Pro, Sol scores 64.6% — ahead of Claude Sonnet 5 (63.2%) and the open-weight GLM-5.2 (62.1%) but well behind Fable 5 (80.3%), which retains the coding-agent crown. On GPQA Diamond, Sol (max) scores 94.1%, matching Gemini 3.1 Pro's 94.3%. On Agents' Last Exam (long-horizon professional workflows), Sol scores 53.6%, Terra 50.4%, Luna 50.3% — all ahead of Claude Fable 5's 40.5% on that benchmark. On Coding Agent Index v1.1, Sol scores 80, above Fable 5's 77.2. OpenAI did not publish GPQA Diamond or MMLU-Pro figures for all three tiers in its launch post.
Benchmarks (published by OpenAI/Vellum):
| Metric | Sol | Terra | Luna | Sol Ultra |
|---|---|---|---|---|
| Terminal-Bench 2.1 | 88.8% | 87.4% | 84.7% | 91.9% |
| SWE-bench Pro | 64.6% | 63.4% | 62.7% | — |
| GPQA Diamond | 94.1% | — | — | — |
| Agents' Last Exam | 53.6% | 50.4% | 50.3% | — |
| Coding Agent Index | 80 | 77.4 | 74.6 | — |
Source: OpenAI — Introducing GPT-5.6 · MarkTechPost — GPT-5.6 launch · Vellum — GPT-5.6 benchmarks explained · Simon Willison — GPT-5.6 family · Engadget — GPT-5.6 goes public · Analytics Vidhya — GPT-5.6 pricing & benchmarks
ChatGPT Work — OpenAI (July 9, 2026)
Who / License: OpenAI; closed product (not a standalone model).
What's notable: Workplace agent built on GPT-5.6 and Codex that can run multi-step tasks for hours, pulling context from connected apps and files to generate reports, spreadsheets, presentations, and websites. Rolled out July 9 to Pro, Enterprise, and Edu users on web and mobile; expanding to Plus and Business over the following days. A direct competitive response to Anthropic's Claude Cowork (launched July 9 as well). Highlights how both OpenAI and Anthropic are converging on long-horizon, app-integrated agents as the primary enterprise product surface.
Source: Forbes — ChatGPT Work launch · Bloomberg — ChatGPT Work · The Next Web
GPT-Live-1 / GPT-Live-1 mini — OpenAI (July 8, 2026)
Who / License: OpenAI; closed, proprietary voice models.
What's notable: Full-duplex voice architecture: unlike prior voice AI, GPT-Live can listen and speak simultaneously, making sub-second interaction decisions (speak, pause, interrupt, call a tool) many times per second. The model handles natural conversational signals (affirmations, filler sounds) natively. For complex queries, it silently delegates to GPT-5.5 in the background and brings the result back into the conversation. GPT-Live-1 is the default voice model for paid ChatGPT users; GPT-Live-1 mini covers free users. No inference API pricing announced at launch. This is the first full-duplex voice model from OpenAI to ship broadly and represents a meaningful architecture shift from the turn-based GPT-4o voice mode.
Source: OpenAI — Introducing GPT-Live · TechCrunch — GPT-Live voice models · MarkTechPost — GPT-Live-1 · The Decoder — full-duplex voice
Open-Weight Front — No New Releases This Week
No new open-weight model releases since the July 8 edition. The standout open models remain GLM-5.2 (Zhipu AI, MIT, 744B MoE, SWE-bench Pro 62.1%, GPQA 91.2%, ~$1.18/M avg) and DeepSeek V4-Pro (MIT, 1.6T MoE, 49B active, $0.44/$0.87 per M tokens, GPQA 90.1%). Both continue to close the gap on closed frontier models at substantially lower cost — GLM-5.2's SWE-bench Pro score is now within 2.5 points of GPT-5.6 Sol.
Head-to-Head — Current Frontier
The table below reflects the state of the frontier as of July 10, 2026. AA Index = Artificial Analysis Intelligence Index v4.1 (162 models evaluated; higher = more capable). Arena Elo figures from the June 2026 leaderboard snapshot; GPT-5.6 entered July 9–10 with ~1,465 (initial, low-vote-count figure). SWE-bench column uses Pro variant throughout — scores on Verified and Pro are not comparable. All prices per 1M tokens (input / output); "—" = not publicly confirmed or not comparable.
| Model | Org | Open? | Arena Elo† | AA Index | GPQA Dia | SWE-bench Pro | Terminal-Bench 2.1 | Context | $/M in/out |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | No | — (new) | 60 (#1) | — | 80.3% | 86.0% | 1M | $10/$50 |
| GPT-5.6 Sol | OpenAI | No | ~1,465 (new) | 59 (#2) | 94.1% | 64.6% | 88.8% / 91.9%‡ | 1M+ | $5/$30 |
| Claude Opus 4.8 | Anthropic | No | ~1,510 | 56 (#5) | — | 69.2% | ~85.0% | 1M | ~$7.22 avg |
| Grok 4.5 | xAI | No | — (new) | 54 (~#4) | — | — | — (beats Opus per xAI) | 500K | $2/$6 |
| Gemini 3.1 Pro Preview | No | ~1,505 | 46 (#13) | 94.3% | 54.2% | — | 1M | $2/$12 | |
| GPT-5.6 Terra | OpenAI | No | — | — | — | 63.4% | 87.4% | 1M+ | $2.50/$15 |
| Claude Sonnet 5 | Anthropic | No | — (new) | — | — | 63.2% | 80.4% | 1M | $2/$10* |
| Qwen 3.7 Max | Alibaba | No (API) | ~1,488 | — | 92.3% | — | — | 1M | $1.25/$3.75 |
| GLM-5.2 | Zhipu/Z.ai | Yes (MIT) | — | 51 (top open) | 91.2% | 62.1% | 81.0% | 1M | ~$1.18 avg |
| DeepSeek V4-Pro | DeepSeek | Yes (MIT) | ~1,410 | — | 90.1% | — | — | 1M | $0.44/$0.87 |
*Claude Sonnet 5 at introductory pricing through August 31, 2026; then $3/$15.
†Arena Elo: models released after mid-June have too few votes for stable rankings; "" = June 2026 snapshot; "— (new)" = model just entered the pool.
‡GPT-5.6 Sol Ultra (4-agent Codex mode) achieves 91.9%; standard Sol is 88.8%.
§Claude Opus 4.8 average pricing from LLM Stats; per-token split not officially confirmed.
§Grok 4.5 AA Index #4 approximate per BuildFastWithAI July 10 coverage; independent benchmark coverage still sparse (only 6 published scores as of July 8).
Reading the table: GPT-5.6 Sol's General Availability today reshapes the competitive picture at the $5/$30 price tier — it matches or beats Gemini 3.1 Pro on GPQA Diamond (94.1% vs. 94.3%) and beats every model except Fable 5 on Terminal-Bench and Coding Agent Index. Claude Fable 5 retains the AA Intelligence Index #1 slot (score 60 vs. Sol's 59) and dominates SWE-bench Pro (80.3% vs. Sol's 64.6%) — still the best agentic coding benchmark available. Terra ($2.50/$15) is notable: 87.4% Terminal-Bench and 63.4% SWE-bench Pro at roughly Sonnet 5 pricing. On the open-weight side, GLM-5.2 remains the standout — its SWE-bench Pro score (62.1%) is within 2.5 points of Sol and less than 1 point below Sonnet 5, at a fraction of the cost, under an MIT license.
Benchmark & Leaderboard Movement
- New AI Intelligence Index leader contender: GPT-5.6 Sol (max) entered the Artificial Analysis Intelligence Index at score 59, just one point behind Claude Fable 5's 60 — the closest any non-Anthropic model has come to Fable 5 since its June launch. Claude Opus 4.8 (max) and GPT-5.6 Sol (high) are both at 56.
- Terminal-Bench 2.1 record extends: Sol Ultra (multi-agent Codex) at 91.9% sets a new state-of-the-art on the benchmark measuring autonomous terminal task completion, overtaking Claude Fable 5 (86.0%) and Grok 4.5's claimed rival. Standard Sol (88.8%) also leads all single-agent figures.
- SWE-bench Pro: Fable 5 moat holds — for now: GPT-5.6 Sol's 64.6% is a solid result, but Claude Fable 5's 80.3% is still ~16 points ahead. This gap is the clearest argument for paying $10/$50 vs. $5/$30 on agentic coding jobs.
- Grok 4.5 hallucination data emerges: Post-launch testing covered by BuildFastWithAI shows a 54% hallucination rate on standard factual tasks, up sharply from Grok 4.3's 25%. Combined with sparse independent benchmark coverage (6 scores across 254 tracked), Grok 4.5's self-reported benchmark claims deserve extra skepticism until third-party evals land.
- Gemini 3.5 Pro delay continues: Reports now cite a target of July 17 for Gemini 3.5 Pro GA, though Google has made no official statement. After slipping from June, then "early July," the model remains in limited enterprise preview. Features cited: 2M token context, Deep Think reasoning layer. Treat any specific score or date as unconfirmed until Google publishes model card and pricing.
- Arena Elo cluster unchanged: The June 2026 top-6 cluster (Claude Opus 4.8 ~1510, GPT-5.5 Pro ~1510, Gemini 3.1 Pro ~1505) is still the most compact top-tier spread on record. GPT-5.6 Sol just entered the pool at ~1465 — expect that score to move significantly as votes accumulate over the next two weeks.
Analysis
For agentic coding and software engineering tasks, the picture this week sharpens: Claude Fable 5 ($10/$50) is the only model above 70% on SWE-bench Pro and retains the AA Intelligence Index #1 slot. But GPT-5.6 Terra ($2.50/$15) at 63.4% SWE-bench Pro is now the most cost-competitive closed-model alternative, and open-weight GLM-5.2 (~$1.18/M avg) is within 1 point of Sonnet 5 on the same benchmark under an MIT license. For terminal-and-agent work specifically, Sol Ultra (91.9% Terminal-Bench, $5/$30 Sol pricing with multi-agent overhead) now leads. For reasoning and knowledge (GPQA Diamond), the top three — GPT-5.6 Sol (94.1%), Gemini 3.1 Pro (94.3%), Claude Mythos Preview (94.6%, invite-only) — are statistically indistinguishable; Gemini at $2/$12 is the accessible pick. For cheap-and-fast, GPT-5.6 Luna ($1/$6) and Claude Sonnet 5 ($2/$10 introductory, through August 31) anchor the value tier; DeepSeek V4-Flash ($0.28/M output, open-weight) is the extreme-budget option. The open-vs-closed gap at the $1–2/M tier has now effectively closed for most non-frontier coding tasks: GLM-5.2 and DeepSeek V4-Pro are genuine alternatives, not catch-up models.
Sources
- OpenAI — Introducing GPT-5.6 (GA announcement)
- OpenAI — Previewing GPT-5.6 Sol (June 26 preview post)
- OpenAI — Introducing GPT-Live
- MarkTechPost — GPT-5.6 three-tier model family launch
- MarkTechPost — GPT-Live-1 full-duplex voice models
- Vellum — GPT-5.6 Sol vs Terra vs Luna benchmarks explained
- Vellum — Claude Sonnet 5 benchmarks explained
- Simon Willison — The new GPT-5.6 family: Luna, Terra, Sol
- Engadget — OpenAI gets permission to roll out GPT-5.6, July 9
- Analytics Vidhya — GPT-5.6 Sol, Terra, Luna pricing and benchmarks
- ExplainX — GPT-5.6 public launch July 9 guide
- DataCamp — GPT-5.6 Sol, Luna, Terra overview
- QCode — GPT-5.6 Sol, Terra, Luna benchmarks and access
- Digital Applied — GPT-5.6 GA pricing, Ultra mode and access
- Forbes — OpenAI launches ChatGPT Work with GPT-5.6
- Bloomberg — ChatGPT Work agent for complex tasks
- TechCrunch — OpenAI GPT-Live new voice models
- The Decoder — ChatGPT listens and talks simultaneously
- BuildFastWithAI — AI News Today July 10, 2026
- BuildFastWithAI — AI News Today July 9, 2026
- BenchLM — Artificial Analysis Intelligence Index scores
- Artificial Analysis — Intelligence Index leaderboard
- Artificial Analysis — LLM leaderboards
- LM Market Cap — GPT-5.6 Sol Pro Arena Elo and rankings
- LocalAI Master — LMArena Leaderboard June 2026 snapshot
- Arena AI — Official leaderboard
- LLM Stats — AI model leaderboard and updates, July 2026
- DataCamp — Claude Sonnet 5 features and pricing
- VentureBeat — GLM-5.2 beats GPT-5.5 on long-horizon coding
- Technology.org — GLM-5.2 coding benchmarks
- Artificial Analysis — DeepSeek V4-Pro and Flash return to open-weights leading pack
- MorphLLM — DeepSeek V4 architecture and pricing
- Crypto Briefing — Google delays Gemini 3.5 Pro to July 2026
- BigGo Finance — Gemini 3.5 Pro launch delayed to July 17
- TechTimes — Gemini 3.5 Pro cleared for July launch
- Presenc AI — Chatbot Arena Elo Leaderboard June 2026
More from News