📊 Full opportunity report: AI Showdown: Qwen3.8-Max's Latest Numbers And The Fable 5 Comparison on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of Qwen3.8-Max, a 2.4-trillion-parameter AI model, with detailed benchmark results. The open weights will ship next week, prompting comparisons with Fable 5 and other top models.

Alibaba has officially released detailed specifications and benchmark results for its Qwen3.8-Max model, confirming it as the second-largest model publicly available with 2.4 trillion parameters. This marks a significant development in the AI industry, as the company also announced that open weights will be available next week, alongside a smaller 27B-parameter checkpoint. The release confirms Alibaba’s position as a major player in large-scale AI model development and sets the stage for direct comparisons with models like Fable 5.

Alibaba’s Qwen3.8-Max features approximately 95 billion active parameters within a 2.4 trillion-parameter sparse mixture-of-experts architecture based on Qwen3.5. The model is multimodal, capable of processing text, images, and videos, with text output. Benchmark results show the model outperforming several competitors on key tests such as Terminal-Bench 2.1 (86.6), PaperBench (93.0), and Parametric CAD Bench (91.5). It demonstrates notable improvements in agentic tasks, with significant jumps in long-horizon reasoning, but trails behind Fable 5 on deep software-engineering benchmarks like SWE-bench Pro and FrontierSWE.

The model was previewed in stealth, initially appearing as an anonymous submission on the Code Arena leaderboard before Alibaba confirmed its identity at the World AI Conference in Shanghai. The company also announced that the open weights, which will be released next week, are intended for local deployment, with a 27B version optimized for single-machine inference. The 2.4T checkpoint is a multi-node artifact, emphasizing its enterprise-scale deployment potential.

At a glance
updateWhen: announced August 3, 2023; benchmark dat…
The developmentAlibaba confirmed the release of Qwen3.8-Max with comprehensive benchmark data, marking a significant milestone in large-scale open-weight AI models.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Industry Impact of Alibaba’s Large-Scale Model Release

The release of Qwen3.8-Max and its benchmark results mark a milestone in large-scale AI development, demonstrating that open-weight models can achieve competitive performance against proprietary models like Fable 5. The detailed benchmark data provides transparency and offers a new reference point for AI researchers and developers. The upcoming open weights could influence the deployment landscape, especially for organizations seeking powerful models that can run on enterprise hardware, potentially democratizing access to large-scale AI capabilities.

However, the model’s strengths in agentic reasoning and multimodal tasks contrast with its weaker performance in certain software-engineering benchmarks, highlighting ongoing challenges in scaling AI for specialized tasks. The distinction between the 2.4T parameters and the 95B active parameters underscores the importance of understanding model efficiency and deployment feasibility, especially as the industry debates the balance between size and performance.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Development and Benchmarking

Alibaba’s AI efforts have been characterized by strategic stealth and selective disclosures, culminating in the recent reveal of Qwen3.8-Max. The model was previewed in July and identified as the mysterious 'kaleb' on the Code Arena leaderboard before official confirmation. The company’s approach has included releasing limited preview endpoints, a detailed benchmark table, and plans for open weights, positioning itself as a key competitor to other large models like GPT-5 and Claude Fable 5.

Historically, Alibaba’s open models have shipped under licenses like Apache 2.0, but the upcoming 2.4T checkpoint’s licensing details remain unpublished. The model’s architecture, based on the Qwen3.5 framework, employs sparse mixture-of-experts to achieve high parameter counts while maintaining efficiency. Prior to this release, industry speculation centered on whether Alibaba could match or surpass the performance of models like Kimi K3 and Moonshot’s 2.8T model, with recent benchmark results affirming its competitive stance.

"We are committed to transparency and open deployment. The upcoming open weights will enable broader access and innovation."

— Alibaba spokesperson

NVIDIA Tesla A100 Ampere 40 GB Graphics Processor Accelerator - PCIe 4.0 x16 - Dual Slot

NVIDIA Tesla A100 Ampere 40 GB Graphics Processor Accelerator - PCIe 4.0 x16 - Dual Slot

  • Memory Capacity: 40 GB GDDR6
  • Host Interface: PCIe 4.0 x16
  • Cooling Type: Passive Cooler

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Licensing and Deployment

Details about the licensing terms for the 2.4-trillion-parameter checkpoint remain unpublished, raising questions about usage rights and commercial deployment. It is also unclear whether the open weights will include full training data, fine-tuning capabilities, or restrictions on commercial use. Additionally, the performance of the 27B variant, optimized for local inference, has not yet been benchmarked publicly, leaving its comparative capabilities uncertain.

Furthermore, the long-term impact of the model’s agentic improvements and multimodal capabilities on real-world applications remains to be seen, as ongoing testing and deployment will reveal more about its practical strengths and limitations.

The Claude AI Advanced Handbook: Model and Effort Economics for Claude Opus 5: Real Cost Per Task, When to Escalate, When to Downgrade, and the Benchmark Rows Anthropic Lost

The Claude AI Advanced Handbook: Model and Effort Economics for Claude Opus 5: Real Cost Per Task, When to Escalate, When to Downgrade, and the Benchmark Rows Anthropic Lost

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Upcoming Release of Open Weights and Industry Response

The open weights for Qwen3.8-Max are scheduled to be released next week, enabling developers and organizations to evaluate and deploy the model independently. Industry observers will closely monitor how the model performs in real-world scenarios, especially in enterprise environments. Alibaba’s next steps may include licensing details, further benchmark disclosures, and potential collaborations or integrations with other AI platforms.

In parallel, competitors will likely respond with their own announcements and benchmark updates, intensifying the ongoing race for large-scale, open-weight AI models. The AI community will also scrutinize the licensing terms and practical deployment options, shaping the future landscape of accessible, high-performance language models.

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

  • Powerful Microcontroller: ESP32-S3 with 16MB Flash and 8MB PSRAM
  • AI Vision & Voice Capabilities: Camera and audio for AI interactions
  • Multiple Development Platforms: Supports Arduino IDE and ESP-IDF

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the key specifications of Alibaba’s Qwen3.8-Max?

It features approximately 2.4 trillion parameters with 95 billion active per query, built on a sparse mixture-of-experts architecture, and is multimodal—handling text, images, and videos.

When will the open weights for Qwen3.8-Max be available?

The open weights are scheduled to ship next week, making the model accessible for local deployment and research.

How does Qwen3.8-Max compare to Fable 5?

In benchmark tests, Qwen3.8-Max outperforms Fable 5 on several metrics like PaperBench (93.0 vs. 80.0) but trails in deep software-engineering benchmarks such as SWE-bench Pro and FrontierSWE.

What are the licensing implications of the open weights?

Details about licensing are still unpublished; historically, Alibaba’s open models have used Apache 2.0, but the upcoming checkpoint’s license remains uncertain, affecting usage rights and commercial deployment.

Source: ThorstenMeyerAI.com

You May Also Like

MiMo Code is now released and open-source

MiMo Code is now publicly available as open-source, enabling broader access and collaboration in wireless communication development.

Apple’s New SpeechAnalyzer API, Benchmarked Against Whisper And Its Predecessor

Apple’s new SpeechAnalyzer API is tested against Whisper and its predecessor, showing promising performance. Details on capabilities and implications inside.

Show HN: Infinite canvas notes in the non-Euclidean Poincaré disk

A new project introduces an infinite, non-Euclidean canvas for note-taking using the Poincaré disk model, enabling unique visual organization.

Different Game, or Already Lost? Reading Mistral’s Sovereignty Bet

Mistral’s Paris summit recast the French AI company as a full-stack provider, putting its sovereignty pitch and compute limits in focus.