Qwen3.8-Max's AI Performance: How Close Is It To The Top?

📊 Full opportunity report: Qwen3.8-Max's AI Performance: How Close Is It To The Top? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has officially released benchmarks for its Qwen3.8-Max model, confirming it has 2.4 trillion parameters and performs well in multimodal tasks. While it ranks high in some benchmarks, it trails in others, especially deep software engineering. The open weights will be available next week, but the model’s full capabilities and licensing details remain unclear. For more on AI performance, see Harnessing GPT-5.6.

Alibaba has officially published benchmark results for its Qwen3.8-Max model, confirming it has 2.4 trillion parameters and ranks among the top multimodal AI models. This development follows two weeks of speculation and stealth previews, culminating in a public release that includes detailed performance metrics and an upcoming open-weight release scheduled for next week.

The Qwen3.8-Max model features 2.4 trillion total parameters, with approximately 95 billion active parameters per query, utilizing sparse mixture-of-experts architecture based on Qwen3.5. It is multimodal, capable of processing text, images, and videos, with text output.

Benchmark results, conducted on Alibaba’s own testing environment, position Qwen3.8-Max at 86.6 on Terminal-Bench 2.1, surpassing models like Claude Opus 4.8 and Fable 5, but trailing behind GPT-5.6 Sol at 88.8. It scored highest in PaperBench at 93.0 and excelled in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5.

However, in deep software engineering benchmarks like SWE-bench Pro and FrontierSWE, it lagged significantly behind Fable 5, with scores of 67.7 and 73.5 respectively, indicating limitations in these areas. Notably, the model demonstrated substantial improvements over its predecessor in agentic execution, tripling scores in some benchmarks, such as DeepSWE from 21.6 to 56.6.

At a glance
reportWhen: announced August 3, 2023; benchmarks re…
The developmentAlibaba announced the official benchmarking and upcoming open release of its Qwen3.8-Max model after weeks of stealth preview and speculation.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmarking and Open-Weight Release

This development signifies Alibaba's move toward transparency and competitiveness in the AI field, with a model that ranks highly in several benchmarks and demonstrates strong multimodal and agentic capabilities. The upcoming open weights will allow wider access, but their size and licensing restrictions suggest limited immediate deployment for individual developers. The benchmarks confirm the model's strengths and weaknesses, influencing future AI research and commercial applications.

Toshiba Canvio Basics 4TB Portable External Hard Drive USB 3.0, Black - HDTB540XK3CA

Toshiba Canvio Basics 4TB Portable External Hard Drive USB 3.0, Black - HDTB540XK3CA

  • Design: Sleek matte finish, smudge-resistant
  • Ease of Use: Plug & Play, no software needed
  • Storage Capacity: Adds 4TB to your devices

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba's AI Model Launch and Benchmarking Strategy

Over the past two weeks, Alibaba's Qwen3.8-Max was shrouded in secrecy, initially known only through a stealth preview and anonymous community sightings. It was publicly confirmed during the World AI Conference in Shanghai on July 19, after weeks of speculation following the release of other large models like Moonshot's Kimi K3 and community discoveries of 'kaleb'—later identified as Qwen3.8-Max.

Alibaba's approach involved staged announcements, limited previews, and a strategic release of benchmark results, culminating in today’s detailed performance table. The model's parameters and capabilities have been carefully disclosed, emphasizing its multimodal and agentic strengths, while limitations in software engineering benchmarks remain evident.

"We are committed to transparency and open access, with open weights to be released next week, enabling broader research and deployment."

— Alibaba spokesperson

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Licensing and Practical Deployment

Details about the licensing terms for the 2.4 trillion parameter weights are still unpublished, raising questions about commercial use and open-source status. Additionally, it remains unclear whether the agentic performance gains will be preserved in the open weights, especially after compression and quantization. The full capabilities of the 27B variant for local deployment are also yet to be demonstrated with benchmark data.

Plaud NotePin & Clip AI Voice Recorder Combo, Voice Recorder with Additional Accessory Replacement Set, Transcribe & Summarize with AI Technology, 64GB Memory

Plaud NotePin & Clip AI Voice Recorder Combo, Voice Recorder with Additional Accessory Replacement Set, Transcribe & Summarize with AI Technology, 64GB Memory

  • Includes Essential Accessories: Magnetic Clip and Pin Attachments included
  • Multilingual Conversation Capture: Supports 112 languages with AI transcription
  • Advanced AI Technology: Utilizes GPT-5.5, Claude Sonnet 4.6, Gemini 3.1 Pro

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba's Qwen3.8-Max and Industry Impact

The open weights are scheduled for release next week, enabling researchers and developers to evaluate the model's real-world performance. Expect further benchmark publications, licensing clarifications, and potential deployment in commercial applications. The AI community will closely monitor whether the agentic improvements are sustained and how Alibaba's model influences the competitive landscape.

DULIWO Model Scriber Tool Kit, 7-Blade Chisel Set for Gunpla

DULIWO Model Scriber Tool Kit, 7-Blade Chisel Set for Gunpla

  • Complete Model Kit Tools: Includes scribe, drill, tweezers, and brush
  • High-Quality Blades: Tungsten steel, wear-resistant, long-lasting sharpness
  • Ergonomic Handle: Lightweight, non-slip aluminum alloy handle

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main strengths of Qwen3.8-Max?

It excels in multimodal tasks, such as processing images and videos alongside text, and demonstrates significant improvements in agentic execution and long-horizon reasoning.

When will the open weights be available?

Alibaba announced that the 2.4 trillion parameter weights will be released next week, though licensing details are still pending.

How does Qwen3.8-Max compare to other large models like GPT-5?

In benchmark tests, it ranks just below GPT-5.6 in some areas, but outperforms models like Claude Fable 5 and Claude Opus 4.8 in several tasks, especially multimodal and agentic benchmarks.

What limitations does Qwen3.8-Max have?

The model underperforms in deep software engineering benchmarks, with scores significantly lower than Fable 5, indicating room for improvement in those areas.

Will the open weights be suitable for local deployment?

The 2.4 trillion parameter checkpoint is too large for single-machine deployment; however, the 27B variant is designed for local use and will likely perform well in practical, smaller-scale applications.

Source: ThorstenMeyerAI.com

You May Also Like

Amazon’s Fire HD 10 tablet just got a refresh with a bit more RAM

Amazon’s Fire HD 10 tablet has been updated with increased RAM from 3GB to 4GB, maintaining other specs but costing $15 more. Available now.

Train sim created by just one person is being called the best ever made

A solo developer’s train simulation game is being hailed as the best ever made, gaining widespread acclaim for its quality and innovation.

Destiny 2 Players, Time Is Running Out To Melt Bosses With This Exploit

Players in Destiny 2 can currently use an exploit to quickly defeat bosses, but the window to do so is closing as Bungie prepares a fix.

Everything Announced At The UploadVR Showcase – Summer 2026

A comprehensive summary of all VR game and experience announcements from the UploadVR Summer 2026 Showcase, including release dates and details.