Alibaba Previews Qwen3.8-Max, a 2.4 Trillion-Parameter Multimodal Model, Days After Moonshot’s Kimi K3 Open-Weight Launch
On July 19, Alibaba’s Qwen team previewed Qwen3.8-Max-Preview , the next flagship in the Qwen family.

On July 19, Alibaba’s Qwen team previewed Qwen3.8-Max-Preview , the next flagship in the Qwen family. The research team describes it as a 2.4 trillion-parameter model, ‘second only to Fable 5’ among the systems it benchmarked. The preview is live now. The benchmark table, model card, and license are not.
The July 19th 2026 announcement landed during the World AI Conference (WAIC) in Shanghai. It also arrived two days after Moonshot AI released Kimi K3 , a 2.8 trillion-parameter open-weight model. The timing is the story as much as the model.
This article separates what Alibaba confirmed from what it only claimed. Every performance figure below carries that caveat.
The Qwen account posted that Qwen3.8 is launching and going open-weight soon. It called the model ‘one of the most powerful available today, comparable to leading frontier systems.
The preview build is real and purchasable. Access runs through Alibaba’s Token Plan subscription. The preview is offered at 10% of standard pricing.
Qwen developer Shuai Bai added technical detail . He described Qwen3.8 as the team’s first multimodal model above 1 trillion parameters. It processes text, images, video, and documents. Alibaba team states the model should beat Qwen3.7-Max on coding, full-stack development, data analysis, and office workflows.
Status as of July 19, 2026. Toggle a fact type using the tabs above.
Total count is not usable compute. A sparse MoE model activates only a fraction of these parameters per token.
Estimate for weights alone; add roughly 20–30% for the KV cache and runtime overhead. Active-parameter serving would be far cheaper, but Alibaba has not published that number.
Qwen3.7-Max, May 2026, closed weights. Alibaba’s historic edge has been price-to-performance, not topping a leaderboard.
Total parameter count is not the same as usable compute. This distinction matters more than the main number. Qwen’s own history proves the point.
Qwen3-235B-A22B carries 235 billion total parameters but activates 22 billion per token. Qwen3-30B-A3B activates roughly 3 billion. Both are sparse MoE designs, and Qwen’s Max tier is too.
For Qwen3.8, the active-parameter count is the number nobody has. Without it, the 2.4T highlighted parameters says little about serving cost. As Startup Fortune calculated , a 2.4T model at 4-bit precision needs roughly 1.2 terabytes for weights alone. A single Nvidia H200 carries 141GB. Even eight cards leave awkward math.
Source: MarkTechPost