AI Blog

Claude Mythos 5 vs GPT-5.6 vs Gemini 3.2 vs Qwen 3.7 vs DeepSeek V4.1: The June 2026 Frontier Refresh

Five frontier-tier models shipped inside a two-week window in June 2026. The differences are no longer about who tops MMLU — each lab is now betting on a different axis: agentic computer use, reasoning cost, multimodal latency, or pure price floor. Pick the axis before you pick the model.

By Agentic AI Wiki 19 min read

Five frontier models, two-week window — and the MMLU leaderboard is no longer the axis worth tracking. As of late June 2026, each lab is betting on a different one, so the procurement question is not which is best: it is pick the axis before you pick the model. Anthropic shipped Claude Mythos 5 GA (with Fable 5 in preview), OpenAI shipped GPT-5.6, Google shipped Gemini 3.2, and the Chinese cluster fired back with Qwen 3.7 and DeepSeek V4.1 — all of it inside the same fortnight.

At a glance

Five flagships, five different axes of attack. The table below sets the basics; the bar chart and feature matrix that follow show where each model actually leans hardest.

Model Released Pricing (USD / 1M out) Headline strength
Claude Mythos 5 (GA) Mid-June 2026 ~$15 Coding & agentic reasoning
GPT-5.6 Mid-June 2026 ~$10 Autonomous computer use
Gemini 3.2 Mid-June 2026 ~$6 Multimodal latency & price/quality
Qwen 3.7 Mid-June 2026 ~$2 Open weights, multilingual
DeepSeek V4.1 Mid-June 2026 ~$1.40 Price floor & reasoning economics

Snapshot as of late June 2026. Pricing tiers shift weekly across providers; context length runs from 200K (Claude, GPT) through 1M (Gemini) to 256K natively for Qwen and 128K for DeepSeek — figures shift between API tiers and weight releases, so confirm against the active model card before committing.

SWE-bench Verified (% success, snapshot) — frontier model comparison Horizontal bar chart: Claude Mythos 5 leads at 72%, GPT-5.6 at 68%, Gemini 3.2 at 61%, Qwen 3.7 at 58%, DeepSeek V4.1 at 55%. All scores are illustrative late-June 2026 snapshots. SWE-bench Verified (% success, snapshot) 0 25 50 75 100 Claude Mythos 5 72% GPT-5.6 68% Gemini 3.2 61% Qwen 3.7 58% DeepSeek V4.1 55%
SWE-bench Verified snapshot, late June 2026. The bars are tight enough that the leaderboard will flip again next quarter — read the order, not the deltas.
Frontier model capability matrix — June 2026 Heatmap comparing five frontier models across five capability axes. Claude Mythos 5: strong reasoning, strong computer use, medium multimodal, medium price/perf, weak open-weights. GPT-5.6: strong across reasoning, computer use, multimodal; medium price/perf; weak open-weights. Gemini 3.2: strong reasoning and multimodal; medium computer use; strong price/perf; weak open-weights. Qwen 3.7: strong reasoning; medium computer use and multimodal; strong price/perf and open-weights. DeepSeek V4.1: strong reasoning; weak computer use; medium multimodal; strong price/perf and open-weights. Frontier model capability matrix Reasoning Computer use Multimodal Price / perf Open- weights Claude Mythos 5 Strong Strong Medium Medium Weak GPT-5.6 Strong Strong Strong Medium Weak Gemini 3.2 Strong Medium Strong Strong Weak Qwen 3.7 Strong Medium Medium Strong Strong DeepSeek V4.1 Strong Weak Medium Strong Strong Weak Medium Strong
Where each model leans hardest. The matrix is jagged on purpose: nobody is strong everywhere, and the strengths almost never overlap.

Release timeline

Five frontier models in one fortnight (June 2026) Horizontal date axis from June 3 to June 21 2026, with a shaded convergence window covering roughly 14 days. Five lab release events are marked: Claude Mythos 5 on June 9, GPT-5.6 on June 11, Gemini 3.2 on June 13, Qwen 3.7 and DeepSeek V4.1 on June 17, and a Chinese price-pressure cluster on June 20. Five frontier models in one fortnight (June 2026) ~14-day convergence window Jun 3 Jun 6 Jun 9 Jun 12 Jun 15 Jun 18 Jun 21 Claude Mythos 5 GA + Fable 5 preview GPT-5.6 OpenAI reasoning refresh Gemini 3.2 Google multimodal upgrade Qwen 3.7 + DeepSeek V4.1 Chinese frontier cluster GLM-6 / Hunyuan / ERNIE Price-pressure response
Five frontier releases inside a 14-day window. This is the thesis of the post: the convergence is the story.

The convergence is not coincidence. Anthropic and OpenAI run roughly quarterly cadences and both had Q2 2026 slots reserved for their respective flagship updates. Google's Gemini cycle had been telegraphed at I/O. What pulled the Chinese cluster into the same window was specifically DeepSeek V4's April 2026 pricing, which compressed the per-token economics for reasoning workloads to a level the other Chinese labs could not ignore. Qwen 3.7 and the V4.1 refresh that immediately followed both ship inside the same fortnight as the Western flagships — that is a deliberate response, not a calendar accident.

The interesting structural fact is that no two of the five models are racing on the same axis. Anthropic and OpenAI are competing for the agentic crown (one leading on coding, the other on autonomous computer use). Google is competing on multimodal price/quality. The Chinese cluster is competing on per-token cost and weights availability. The two-week window is the moment all five bets became legible at once.

Claude Mythos 5 GA (+ Fable 5 preview)

Mythos 5 went GA in mid-June 2026 as the production sibling of a two-line strategy: Mythos for production, Fable for research preview. The split is procurement-friendly by design. Mythos carries the safety-evaluation paperwork an enterprise buyer expects — vulnerability-aware reasoning evals, model-card disclosure, the usual SOC-shaped diligence — and ships with the same stable API surface as the Claude 4 family it replaces. Fable 5 ships in preview alongside it, on the same model architecture but with research-grade post-training extensions (longer thinking traces, experimental tool patterns, sharper exploratory reasoning). The names matter: nobody has to ask whether a model is procurement-ready, because the line tells you.

On the snapshots that are currently visible, Mythos 5 leads on the two axes Anthropic has been winning all year — coding and agentic reasoning. On SWE-bench Verified it lands at the top of this five-model set; on long-horizon coding tasks the lead is wider than the benchmark gap suggests. The computer-use story carries over from the Vercept integration covered in the previous post on computer-use offerings, which means Mythos 5 inherits the OSWorld lead rather than re-litigating it.

Where Mythos 5 does not lead is price. At roughly $15 per million output tokens it is the most expensive of the five by a comfortable margin, and the gap to the Chinese cluster is an order of magnitude. The pitch is unchanged from the Claude 4 era: if your unit economics survive the per-token cost, the quality and the procurement story are the cleanest in the set. If they do not, you read on. For teams using Anthropic primarily as the orchestrator across other components, see agent frameworks for the integration shapes that make the cost math work.

GPT-5.6

GPT-5.6 is OpenAI on schedule. It is not a discontinuous jump — it is the on-cadence point release that consolidates everything OpenAI shipped between GPT-5 and now into one stable surface. Reasoning quality nudges up; tool-use trajectories get tighter; the model card lists the usual incremental wins on math, code, and long-context recall. The headline story is not the eval numbers. It is the depth of the integration with the OpenAI agent stack.

Specifically, GPT-5.6 is the first OpenAI flagship that ships with Codex Background CU and Operator as first-class targets rather than as adjacent products. The same model drives the cloud Chromium in Operator, the long-running development tasks in Codex Background CU, and the in-conversation "Agent" surface in ChatGPT — without per-surface fine-tunes that drift apart between releases. That unification is the autonomy lead: when the agent has to cross from one tool to another inside a task, the model does not change underneath it. For the architectural shape of those surfaces and how they compare to the other computer-use bets, see Post 1, Claude Computer Use vs Codex CU vs Operator vs Gemini.

At roughly $10 per million output tokens GPT-5.6 prices below Mythos 5 but well above Gemini and the Chinese cluster. The position is coherent: if you are buying the integrated autonomy story, you accept the price; if you are buying a model to call from your own orchestrator, the per-token line is the one to compare. Headline coding scores trail Mythos 5 by a few points on the visible benchmarks — small enough that a release or two from now the order could flip again, and small enough that the autonomy delta matters more than the percentage.

Gemini 3.2

Gemini 3.2 is the mid-cycle point release in Google's 3.x line — not a generational change, but a refresh that sharpens the parts of the model where Google has the structural advantage. Multimodal latency is the headline: video understanding, image reasoning, and screenshot grounding all improved measurably over 3.1, and the round-trip on the 1M-token context window is the lowest of any model that supports that length. For products where the agent has to look at a frame, decide, and act before the user notices a pause, Gemini 3.2 is the one whose architecture is shaped for it.

The Workspace integration is the second piece of the bet. Gmail, Calendar, Docs, Sheets, and Drive all get first-party Gemini surfaces, and the 3.2 update tightens those flows — summarization, draft replies, sheet formula synthesis — without changing the model API. The same model drives the consumer Gemini app, the Workspace surfaces, and the Vertex AI tier for enterprise; the trade is that the deepest integration only pays off if your users are already inside the Google ecosystem.

Pricing at roughly $6 per million output tokens puts Gemini 3.2 in the middle of the five — cheaper than the Western pair, more expensive than the Chinese cluster — and the price/quality position is the one Google has been steadily refining. On general reasoning the model trails Mythos 5 and GPT-5.6 on the current snapshots; on multimodal and on price-per-quality for the workloads it targets, it leads. The 1M-token context is wide enough that long-document RAG starts to look like just "fit it in the window" rather than a retrieval problem — useful, with the usual caveat that long-context recall degrades faster than the headline number suggests.

The Chinese cluster

The two leaders inside the Chinese cluster are Qwen 3.7 and DeepSeek V4.1, and they are the reason the late-June window included five flagships rather than three. Qwen 3.7 extends the Qwen3 family's bet on language coverage and hybrid reasoning with sharper agent tool-use post-training, broader multimodal coverage, and an Apache-2.0 license inherited from the rest of the family. DeepSeek V4.1 follows V4's April pricing salvo with a quality bump — same MLA + sparse-attention architecture from the V3 lineage, refined post-training, and the same per-token economics the V4 line forced into the market.

The cluster behind them sweeps wide. Zhipu's GLM-6 ships an open-weights MoE with hybrid reasoning that competes with Qwen on the multilingual axis. Tencent's Hunyuan Large 3 takes the closed-source Chinese flagship slot with strong CN-domain post-training. Baidu's ERNIE 5.1 lands inside the ERNIE 5 lineage with the same enterprise-Chinese pitch ERNIE has carried for years. ByteDance's Doubao Pro keeps the consumer-scale price floor, optimized for Doubao's deployment context inside the ByteDance product surface. None of these four leads on a frontier benchmark, but together they make the cluster too big to ignore on price and too deep to ignore on weights availability.

The pricing pressure is the story. DeepSeek V4's April 2026 release dropped output token pricing to a level the rest of the Chinese cluster could not match without responding inside weeks — and they did. Qwen 3.7's hosted DashScope tier sits at roughly $2 per million output tokens; DeepSeek V4.1's own API floors at roughly $1.40; the rest of the cluster lands inside that range. Compared to the Western trio at $6, $10, and $15, the cluster is an order of magnitude cheaper for the high-volume reasoning workloads it targets. The compression is not a temporary promotion — it is the price floor the cluster is committed to defending.

Weights availability is the second axis where the cluster pulls ahead. Qwen 3.7 ships under Apache 2.0 with weights on Hugging Face and ModelScope, DeepSeek V4.1 ships under MIT on the same channels, GLM-6 ships open-weights from Zhipu's GitHub. For teams that need to self-host — for data-boundary reasons, for offline deployment, or for the unit-economics math that hosted APIs cannot beat at scale — the Chinese cluster is the only one of the five labs offering a frontier-tier checkpoint to download. The Western trio is closed-API only.

Cross-cutting comparison

Reasoning

On the visible reasoning snapshots — SWE-bench Verified, GPQA, the long-horizon agent traces — Mythos 5 leads narrowly, GPT-5.6 trails by a small margin, and Gemini 3.2 sits a tier below the Western pair on pure reasoning while leading on multimodal-bound reasoning tasks. Qwen 3.7 and DeepSeek V4.1 land in the next tier, with Qwen pulling ahead on multilingual reasoning and DeepSeek pulling ahead on cost-adjusted reasoning (the same answer for a fraction of the spend). The gaps between adjacent ranks are small enough that a release or two from now the order could flip; the gap between the top two and the bottom two is wider but closing each quarter. Read these as snapshots, not as durable orderings — see reading benchmarks for the trap of treating leaderboard order as a stable signal.

Computer use

GPT-5.6 is the autonomous computer-use leader at the API level — Codex Background CU and Operator both ship with the model as a first-class target, and the unified-model story (one checkpoint across all three OpenAI agent surfaces) is the autonomy delta the other four cannot match without a separate fine-tune. Mythos 5 inherits the Claude Computer Use post-Vercept architecture and leads on raw OSWorld score, which means it wins on per-task success when you build the harness yourself but trails on integrated agent products you buy off the shelf. Gemini 3.2 holds the browser-tab-inside-Workspace lane and is the fastest of the three for what it targets. Qwen 3.7 and DeepSeek V4.1 do not ship a first-party computer-use product; the open-weights computer-use stack (browser-use, UI-TARS, CogAgent) runs against them, but the gap to the closed leaders is real and the integration burden lands on you.

Multimodal

Gemini 3.2 leads on multimodal across the board — video, image, audio, screenshot — and the 1M-token context is wide enough that video reasoning starts to look like just a long prompt. Mythos 5 and GPT-5.6 ship strong vision-in capabilities (image reasoning, screenshot grounding for their respective computer-use stacks) without trying to match Gemini's video depth. Qwen 3.7 ships a Qwen3-VL sibling with respectable image reasoning under the same Apache 2.0 license, which makes it the only open-weights flagship in the set with serious vision-in capability. DeepSeek V4.1 stays text-only at the flagship tier; the DeepSeek-VL line is a separate sibling and trails the others on raw vision benchmarks. If your product is multimodal-first, the order is Gemini → Qwen-VL → Mythos/GPT → DeepSeek.

Price/perf

Price per 1M output tokens (USD, snapshot) Horizontal bar chart showing output token pricing. Claude Mythos 5 is most expensive at $15.00 per million tokens, followed by GPT-5.6 at $10.00, Gemini 3.2 at $6.00. The Chinese cluster leads on price: Qwen 3.7 at $2.00 and DeepSeek V4.1 at $1.40 — both shown in accent fill to mark the price-floor thesis. All figures are illustrative June 2026 snapshots. Price per 1M output tokens (USD, snapshot) $0 $4 $8 $12 $15 Claude Mythos 5 $15.00 GPT-5.6 $10.00 Gemini 3.2 $6.00 Qwen 3.7 $2.00 DeepSeek V4.1 $1.40
Snapshot of hosted-API output-token pricing in late June 2026. The chart spans an order of magnitude — Western pair on top, Chinese cluster on the floor.

The price chart is the cleanest of the comparisons because the units are objective and the spread is an order of magnitude. Mythos 5 at $15 and GPT-5.6 at $10 are the premium tier, paid for quality and integration; Gemini 3.2 at $6 is the middle bet, paid for multimodal and Workspace fit; Qwen 3.7 at $2 and DeepSeek V4.1 at $1.40 are the cost floor, paid for raw token economics. The crossover point that decides which tier you can afford is workload-dependent — see cost, quality, and latency for the framing — but for high-volume agent loops that replay long prefixes on every tool call, the order-of-magnitude gap to the Chinese cluster moves unit economics in ways quality differences do not. Snapshot prices, snapshot tiers; what is durable is the relative ordering, not the digits.

Weights availability

The Western trio is closed-API only — Mythos 5, GPT-5.6, and Gemini 3.2 do not ship weights, and there is no public roadmap to do so. The Chinese cluster is the inverse: Qwen 3.7 under Apache 2.0, DeepSeek V4.1 under MIT, GLM-6 and the rest of the cluster open under permissive or near-permissive licenses (a 14B Qwen 3.7 dense sibling runs on a single H100; a 4-bit quantization of DeepSeek V4.1 fits on an 8×H200 node). For procurement contexts that require weights — air-gapped deployments, data-boundary regimes that cannot route through a vendor API, regulated industries with hard offline requirements — the choice is structural: the cluster is the only option. The trade is the geopolitical and audit overhead of running a Chinese-lab checkpoint inside a Western enterprise; see open vs closed models for the framing that makes this trade-off explicit.

When to pick which

Use case Best Western option Best Chinese-cluster option Rationale
Enterprise procurement Mythos 5 (with Fable 5 as the research-preview sibling) Qwen 3.7 hosted via DashScope or a regional partner Mythos's safety paperwork plus the sibling-line split clears procurement cleanly; Qwen's Apache 2.0 is the permissive-license backstop when the buyer needs weights.
Agentic terminal work GPT-5.6 via Codex Background CU Qwen 3.7 + open-source agent loop (Qwen-Agent SDK) GPT-5.6's unified-model autonomy is the off-the-shelf win; Qwen-Agent ships first-party tool patterns when the work must stay on your infra.
Multimodal product Gemini 3.2 — leads on video, audio, and 1M-token context Qwen 3.7 (with Qwen-VL sibling for vision-in) Gemini owns multimodal latency and Workspace integration; Qwen-VL is the only Apache-2.0 vision-language flagship in the set.
Self-host / weights required None — Western trio is closed-API only. Qwen 3.7 (Apache 2.0) or DeepSeek V4.1 (MIT) The choice is structural: only the Chinese cluster ships frontier weights. Pick Qwen for the cleanest license, DeepSeek for the cheapest inference economics.
Price-optimized batch reasoning Gemini 3.2 if Workspace-shaped; otherwise step out to the cluster. DeepSeek V4.1 — the price floor of the set. For batch workloads where the per-token cost dominates the quality delta, DeepSeek's $1.40 tier is the unit-economics answer the Western trio does not have.

FAQ

Is Mythos 5 just Opus renamed?

No. Mythos 5 is the production line of a new two-line naming strategy that supersedes the Claude 4 family — Mythos for GA, Fable for research preview — and it ships with new training and post-training, not a rebadge. The names are the procurement signal: if you see "Mythos" you can deploy it in production with the usual diligence; if you see "Fable" it is research-grade and changes more often. The API surface is compatible with the Claude 4 endpoints so migration is mechanical, but the model behind the endpoint is new.

Which is cheapest for high-volume reasoning?

DeepSeek V4.1, by roughly an order of magnitude versus the Western trio — its hosted API floors at around $1.40 per million output tokens as of late June 2026, against Gemini's $6, GPT-5.6's $10, and Mythos 5's $15. Qwen 3.7 at $2 is close behind, and you can run either Chinese-cluster model on your own infra to drop the cost further. The caveat is the usual one: prices move weekly, provider tiers shift, and reserved-capacity contracts on the Western trio can narrow the gap for committed buyers.

Can I run Qwen 3.7 on a single H100?

Not the flagship MoE checkpoint, but yes — for the dense siblings. The Qwen 3.7 family ships dense models from small (a few B) up to roughly 32B alongside the MoE flagship, all under the same Apache 2.0 license; a 14B or 32B dense Qwen 3.7 runs comfortably on a single H100 with strong tool-use behaviour through the Qwen-Agent SDK. The flagship MoE wants a multi-GPU node. The dense ladder is the part of "open-weights frontier" that the headline number alone hides.

Does GPT-5.6 dethrone Claude on coding?

Not on the visible snapshots. Mythos 5 still leads SWE-bench Verified by a few points and the lead is wider on long-horizon coding tasks where the agent has to manage state across many turns. GPT-5.6's coding score moved up over GPT-5, and the gap to Mythos 5 is small enough that another release could flip it — but as of late June 2026, on the benchmarks visible at draft time, Mythos 5 holds the coding crown. The interesting question is not who leads on the bench but who leads on the loop you actually run.

Are the Chinese weights actually permissively licensed?

Yes, for the two leaders. Qwen 3.7 ships under Apache 2.0, DeepSeek V4.1 ships under MIT — both are unrestricted commercial-use licenses with no MAU cap and no name-derivative restriction. The license text is what it says it is. The practical question for Western buyers is not the license; it is the procurement overhead of running a Chinese-lab checkpoint inside a regulated enterprise — supply-chain provenance, model-card audit, the geopolitical conversation with security. The license is permissive; the deployment conversation is longer.

What's the difference between Mythos and Fable?

Mythos is the production line: GA-stable, safety-evaluated, the same API contract you can sign a procurement document against. Fable is the research-preview line: same model architecture and family, but with experimental post-training extensions — longer thinking traces, sharper exploratory reasoning, prototype tool patterns — and a faster change cadence that is not procurement-friendly. The split is deliberate: nobody has to ask whether a given checkpoint is enterprise-ready, because the line name tells you. Use Mythos for production, Fable for exploration.

Further reading

On this wiki:

  • Llama 4 vs DeepSeek V3 vs Qwen3 vs Mistral Large 3 — the previous frontier post on open-weights flagships; this one refreshes the picture across both open and closed weights as the cluster expanded into mid-2026.
  • Claude Computer Use vs Codex CU vs Operator vs Gemini — the architectural shape of the computer-use surfaces that GPT-5.6's autonomy lead and Mythos 5's OSWorld inheritance both build on.
  • Choosing a model — the constraint-first selection guide that turns "best model" into "best for this constraint" and makes the axes-not-leaders framing actionable.
  • Reading benchmarks — how to read MMLU / GPQA / SWE-bench numbers without getting played by snapshot ordering that flips every release.
  • Open vs closed models — the trade-off framing for the structural split between the Western trio and the Chinese cluster.
  • Cost, quality, and latency — the triangle that decides which tier's per-token economics actually matters for your workload.

Project sources: