Alibaba launched Qwen3.8-Max on August 3, 2026, after a two-week preview that went live on July 19. At 2.4 trillion total parameters with ~95 billion activated per token, it is the largest Qwen model ever shipped — and, critically, the first time Alibaba has put a Qwen-Max-class flagship's weights on HuggingFace and ModelScope. It is also the second 2T+ open model in two weeks, following Moonshot's Kimi K3 (2.8T) by a month.
The positioning is explicit: coding and professional work. Alibaba's own framing is "A New Bar for Coding and Cowork," and the numbers back it up — PaperBench 93.0 beats Claude Fable 5 (88.8) and GPT-5.6 Sol (90.5); Terminal Bench 2.1 lands at 86.6, within two points of Sol; Vision Arena ranks it #2 globally. Pricing is $2 input / $6 output per million tokens, with explicit cache read at $0.17 — roughly 3× cheaper than Kimi K3 on output and 5× cheaper than Fable 5 across the board.
What's New

2.4T total, ~95B active, hybrid attention. Qwen3.8-Max is a sparse MoE with a hybrid full-and-linear attention architecture, supporting up to 1M tokens. It is the second-largest open model ever released (behind only Kimi K3's 2.8T) and the largest Qwen by a wide margin.
First open Qwen-Max-class weights. Alibaba released the weights as Qwen3.8-2.4T-A95B on HuggingFace and ModelScope. Until August, "Qwen-Max" meant API-only; now the flagship tier is downloadable. NVIDIA published a deployment guide the same week: 4,000+ tokens/second per GPU and 350+ tokens/second per user on an 8×GB300 NVL72 node.
Native multimodality across the full loop. Text, image, and video are understood by one backbone — but the headline is not input. Vision is a feedback loop: Qwen3.8-Max inspects its own rendered output, detects misaligned UI or wrong object orientation, revises its plan, and corrects. It pairs coding with GUI operation ("Hybrid Agent") and was benchmarked on RecreationBench, a black-box app-recreation task across Ubuntu, macOS, Windows, Android, and web.
Arena position: 5th Text, 2nd Vision. CGTN's launch coverage put it "among the world's leading models," trailing only Anthropic's Claude family on the overall Arena leaderboard. On CodeArena (frontend web) it placed 3rd globally.
Pricing at $2/$6 with $0.17 cache. QwenCloud lists input at $2/M, output at $6/M, implicit cache read at $0.25/M, explicit cache read at $0.17/M. Explicit cache write is $2.50/M. Alibaba Cloud's international list is $1.65/$4.95 with context caching.
Architecture

| Component | Qwen3.8-Max | Qwen3.7-Max | Kimi K3 |
|---|---|---|---|
| Total params | 2.4T | ~1T class | 2.8T |
| Active params | ~95B | — | 104B |
| Architecture | Sparse MoE, hybrid full+linear attention | MoE | Stable LatentMoE, KDA+Gated MLA |
| Context | 1M tokens | 1M | 1M |
| Modality | Text + image + video (native) | Text + vision | Text + image + video (native) |
| Open weights | Yes (Qwen3.8-2.4T-A95B) | API-only | Yes (July 27) |
| Input price | $2/M | — | $3/M |
| Output price | $6/M | — | $15/M |
| Cache read | $0.17–0.25/M | — | $0.30/M |
| Served throughput | 4,000+ tok/s/GPU (GB300) | — | ~37 tok/s (AA measured) |
Hybrid attention, configurable reasoning. The model interleaves full attention with linear attention for long-context efficiency, and exposes configurable reasoning effort (Non-Thinking and Thinking modes). NVIDIA's deployment notes 4,000+ tok/s per GPU on GB300 NVL72 — a number that matters because it is roughly an order of magnitude higher than the ~37 tok/s Artificial Analysis measured for Kimi K3 on comparable hardware.
Qwen-MM-Plugins for multimodal agents. Alibaba also shipped Qwen-MM-Plugins, a harness extension library giving any agent framework image/video processing, multimodal memory, dynamic resolution, and visual tool use — plus special adapters for video editing, Blender, and CAD. The point is that Qwen3.8-Max is not just a model; it is a reference stack for building multimodal-native agents.
Harness portability as a feature. Qwen says the model delivers comparable performance across QwenWork, Claude Code, Codex, OpenClaw, and Hermes. That is a deliberate design goal — builders should not have to rewrite their scaffold to use it.
Benchmarks

From Alibaba's official model card (all best-of-harness, max reasoning effort):
| Benchmark | Qwen3.8-Max | Opus 4.8 | Fable 5 | GPT-5.6 Sol | Qwen3.7-Max |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 86.6 | 84.6 | 84.6 | 88.8 | 74.5 |
| DeepSWE 1.1 | 56.6 | 59.0 | 70.0 | 73.0 | 21.6 |
| PaperBench | 93.0 | 80.3 | 88.8 | 90.5 | 64.8 |
| AndroidBench | 75.1 | 69.8 | 84.5 | 74.0 | 56.5 |
| CoWorkBench | 74.8 | 72.3 | 75.9 | 71.5 | 64.6 |
| JobBench | 53.4 | 48.4 | 57.4 | 45.4 | 31.3 |
| SkillsBench | 70.2 | 65.1 | 70.9 | 73.5 | 61.2 |
| GPQA Diamond | 92.6 | — | — | 94.1 | — |
| VideoMME (w/ sub) | 90.4 | 85.4 | — | 89.5 | 88.0 |
| VideoMMU | 88.7 | 75.3 | 81.2 | 85.0 | 85.4 |
| MLVU (M-Avg) | 90.8 | 53.4 | — | 87.6 | 87.4 |
The story in that table is coding and cowork at frontier-adjacent quality. PaperBench is the standout — 93.0 beats both Fable 5 and Sol, and it nearly doubles Qwen3.7-Max (64.8). Terminal Bench 2.1 at 86.6 is within 2.2 of Sol. On video benchmarks it leads every listed competitor (VideoMME 90.4, VideoMMU 88.7, MLVU 90.8).
The honest gaps are long-horizon software engineering: DeepSWE 1.1 at 56.6 trails Fable 5 (70.0) and Sol (73.0), and Qwen itself notes QwenSWEBench (80.7) and QwenQoderBench (58.4) are still behind Fable 5. The 16-day autonomous coding demo — Qwen3.8-Max built "oh-my-cli", a self-evolving agent framework, from scratch over 16 days without human intervention — is the counter to that gap, and it is a real project, not a benchmark number.
What It Means for Developers

The open-weight frontier now has two serious options. Kimi K3 (2.8T, $3/$15, #1 frontend code, slower at 37 tok/s) and Qwen3.8-Max (2.4T, $2/$6, #2 vision, much higher throughput on NVIDIA hardware). Pick K3 for frontend and marathon coding; pick Qwen3.8 for video understanding, cowork, and throughput-sensitive serving.
Price-to-performance moved again. At $2/$6 with $0.17 explicit cache read, Qwen3.8-Max is the cheapest frontier-class open-weight option on the market. Cache-read pricing is 12× cheaper than fresh input — for RAG and agent loops with repeated context, architect for explicit cache and the effective cost collapses.
The Hermes Agent story matters for agent builders. QwenCloud ships Hermes Agent (the Nous Research open-source terminal coding tool, MIT licensed) against Qwen3.8-Max. And the model itself demonstrated 16-day autonomous coding: it built a self-evolving agent framework — establishing an engineering loop with user feedback, community practices, and self-test data, iterating through code generation, testing, previewing, and log analysis. That is the bar for "agent that ships itself" in August 2026.
Vision is now a feedback loop, not an input. If you are building UI generation, design-to-code, or GUI agents, Qwen3.8-Max's ability to render its own output, inspect it, and self-correct is the difference between a demo and a product. Pair it with Qwen-MM-Plugins for video editing, Blender, and CAD workflows.
Self-hosting is now realistic at the GB300 tier. NVIDIA's reference deployment on 8×GB300 NVL72 hits 4,000+ tok/s per GPU. If you have that hardware, the open weights are usable in production — not just downloadable as a museum piece.
vs DeepSeek V4 and Kimi K3 at a glance. DeepSeek V4-Pro (1.6T/49B active) is the quality/throughput middle child; V4-Flash (284B/13B) is the $0.14/M batch tier. Kimi K3 (2.8T/104B) leads frontend and marathon coding but is slow and expensive per output token. Qwen3.8-Max (2.4T/95B) is the vision + cowork + throughput option. The 2026 open-weight stack now has three real lanes, and you route across them by task.
Bottom Line
Qwen3.8-Max is Alibaba's strongest model and the first Qwen-Max-class release with open weights. It does not top every leaderboard — Fable 5 and GPT-5.6 Sol still lead on hard long-horizon coding — but it beats both on PaperBench, ranks 2nd on Vision Arena, leads every listed video benchmark, and prices at $2/$6 with $0.17 cache read. For builders who want open weights at frontier-adjacent quality with a vision feedback loop and NVIDIA-grade throughput, it is the default choice of August 2026.
Resources
- Qwen3.8-Max Official Blog — release post, full benchmark table
- QwenCloud — Qwen3.8-Max — API access, $2/$6 per M tokens
- Alibaba Cloud — Qwen3.8-Max Weights Release — Qwen3.8-2.4T-A95B on HuggingFace/ModelScope
- NVIDIA — Serving Qwen3.8-2.4T-A95B on GB300 NVL72 — 4,000+ tok/s per GPU
- CGTN — Alibaba unveils Qwen3.8-Max — launch coverage
Qwen3.8-Max is available via API at qwencloud.com under the model ID qwen3.8-max, with open weights released as Qwen3.8-2.4T-A95B on HuggingFace and ModelScope.