Back to feed
News Story
APriority71
Latent Space
1 sources

Qwen 3.8 Max (2.4T) and 27B: New Open Weights Models for Coding and Cowork

Qwen 3.8 Max, a massive 2.4T parameter model, and a 27B model have been announced with open weights promised, marking a significant return to open model releases after management changes. The model demonstrates breakthrough capabilities in autonomous coding, research, and hardware design, positioning it as a top open model despite competition from Kimi K3.

SynthePulse Insight · AI deep reading

Qwen 3.8 Max: The Return of the 2.4T Open-Source Giant and the Competitive Landscape

Version 1 · 1 source

After management changes and a pivot to closed-source APIs, Alibaba's Qwen makes a strong return to the open-source frontier with Qwen 3.8 Max, a 2.4T-parameter model. The company promises to release the weights next week and positions the model against top closed-source rivals in coding, agentic tasks, and multimodal domains. This article, based on release information and third-party evaluations, outlines its capabilities, pricing, and market positioning.

  • Qwen 3.8 Max is a 2.4T-parameter flagship model with API pricing of $2/$6 per million tokens, and the company promises to open-source it next week, along with the 27B version.
  • Official claims include 10+ days of autonomous coding, 500+ rounds of chip design optimization, 365 days of e-commerce strategy, and native multimodal visual feedback.
  • Third-party evaluations show it ranks 4th on Frontend Code Arena (Elo 1668), 2nd on Vision Arena, and 2nd among open-source models on Vals Index.
  • Vals Index score of 66.1 ties with Claude Opus 4.7, but testing costs are about 2.3x lower ($2.68 vs $6.17).
  • SWE-bench score of 87.3% leads GPT-5.5 and GLM-5.2, but trails Claude Opus 4.8 (89.2%).
  • Compared to Qwen 3.7 Max, Vals Index improved by 8.6 points (57.5 to 66.1) in about 2.5 months, with a price reduction.
Open section navigationRelease Background and Core Specifications

Release Background and Core Specifications

After last year's 'Qwen Exodus' and management's shift toward closed-source APIs, outsiders questioned whether this leading open-source lab would continue releasing relevant models. On August 3, 2026, Alibaba Qwen announced Qwen3.8-Max as its 'strongest model to date,' promising to open-source the weights next week, along with Qwen3.8-27B.

Official key specifications include: 2.4T total parameters, 1M token context window, API pricing of $2/$6 per million input/output tokens (cached tokens $0.25), and support for low/medium/high reasoning effort modes, compatible with OpenAI and Anthropic protocols. Third-party summaries report approximately 95B active parameters (MoE activation rate of about 4%).

The model is positioned for coding, long-horizon agentic work, and multimodal reasoning, with official claims of 10+ days of autonomous coding, 500+ rounds of chip design optimization, and 365 days of e-commerce strategy execution, emphasizing vision as part of the execution loop rather than just input.

Officially Claimed Capability Highlights

The official announcement lists several breakthrough capabilities: in autonomous coding, it builds a self-evolving coding framework and runs multi-week unattended operations; in autonomous research, it reconstructs a paper pipeline and runs a 125-hour iterative loop, inventing a new data selection method that surpasses the original paper's baseline by +2.71 points.

In a data science competition, it competed against 526 human teams and entered the top 13% within 24 hours (surpassing 87% of human teams); in chip design, it completed the full flow from RTL to physical layout, reducing gate count from 8298 to 678, a 81% area reduction, and meeting timing closure at 500MHz.

In an e-commerce simulation benchmark (365 days of operation), it achieved a 4.16x return (balance ¥416,252) through game-theoretic negotiation and inventory planning; it also released Qwen-MM-Plugins to extend multimodal capabilities to existing agent frameworks.

Third-Party Evaluations and Rankings

Independent evaluations show strong performance: Frontend Code Arena ranks 4th (Elo 1668), behind only Claude Opus 5 Max (1705) and Kimi K3 Max (1676), and roughly tied with Claude Opus 5 High (1669); Vision Arena ranks 2nd (1305), just 13 points behind Claude Fable 5 High.

On Vals Index, Qwen3.8-Max ranks 2nd among open-source models with a total score of 66.1, tying with Claude Opus 4.7, but with testing costs about 2.3x lower ($2.68 vs $6.17). SWE-bench score is 87.3%, leading GPT-5.5 (82.6%) and GLM-5.2 (83.3%), but trailing Claude Opus 4.8 (89.2%).

Vals also reports progress pace: Qwen 3.7 Max scored 57.5, while 3.8 Max scored 66.1, an improvement of 8.6 points in about 2.5 months, with prices dropping from $2.50/$7.50 to $2.00/$6.00. Additionally, some users claim Qwen 3.8 surpasses Fable 5 on Terminal Bench, but this is a personal opinion.

Market Positioning and Competitive Landscape

These data place Qwen3.8-Max in the same deployment category as giant sparse models like Kimi K3 and GLM-5.2, rather than the local practical tier of 30B-70B. Observers believe this marks a direct competition between China's open-source frontier and top Western closed-source models, especially in coding, agentic workflows, and multimodal tasks.

At the same time, multiple infrastructure and application builders (such as Baseten, Hermes Agent, Command Code) have confirmed support or integration plans. Some users visualize benchmark differences and believe 'Opus 4.8 is largely covered by 3.8-Max,' but this is a personal inference.

It is important to note that officially claimed capabilities (such as 10+ days of autonomous coding, 500+ rounds of chip optimization) are mostly vendor-reported and lack independent verification; while third-party evaluations are strong, some data come from summary tweets (e.g., ZhihuFrontier) and should be treated with caution.

Credibility boundary

This report is based on Latent Space's summary coverage, which includes official announcements, third-party evaluations, and personal opinions. Official capability claims (such as days of autonomous coding, rounds of chip design optimization) are vendor-reported and not independently verified; third-party evaluation data (such as Vals Index, SWE-bench) come from institutions like Vals AI, but some details (such as active parameters, benchmark scores) are derived from third-party tweet summaries and require further verification.

Insight takeaway

Qwen 3.8 Max, with its 2.4T parameters and open-source promise, demonstrates competitive strength against top closed-source models in coding, agentic tasks, and multimodal domains, but the extreme capabilities claimed officially still need independent verification. Its rapid iteration and price reduction strategy may intensify competition between open and closed-source models and boost the global influence of China's open-source frontier.

Primary report

Latent Space

Primary source