Back to feed
News Story
APriority85
Artificial Analysis (X)
1 sources

Alibaba Releases Qwen3.8 Max, Plans Open-Source Weights

Alibaba has released Qwen3.8 Max, a 2.4T-parameter MoE model scoring 56 on the Artificial Analysis Intelligence Index, matching Claude Opus 4.8 but trailing Kimi K3. The company plans to open-source the weights next week, marking a strategic shift and making it the largest open-weight model from Alibaba.

SynthePulse Insight · AI deep reading

Qwen3.8 Max Released: Performance Leap but Cost and Hallucination Concerns Linger

Version 1 · 1 source

Alibaba releases Qwen3.8 Max, scoring 56 on the Artificial Analysis Intelligence Index, tying Claude Opus 4.8, but costs double, hallucination rates rise, and the shift in open-weight strategy draws attention.

  • Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, 10 points higher than its predecessor Qwen3.7 Max, tying Claude Opus 4.8, but trailing open-weight leader Kimi K3 by 1 point.
  • The model has 2.4T total parameters, 95B active, 1M context, supports text, image, and video input, but weights are planned for release next week; if delivered, it will be Alibaba's largest open-source model.
  • Cost per task is $1.14, more than double the previous generation, mainly due to more rounds in agentic evaluations, not higher token prices.
  • Score drops 10 points on AA-Omniscience, hallucination rate rises from 23% to 40%, showing the model tends to attempt rather than abstain when unable to answer.
  • GDPval-AA Elo score of 1739 surpasses Kimi K3 but trails Claude Opus 5, with some gains attributed to increased rounds per task.
Open section navigationPerformance Leap: Intelligence Index Ties Top Models

Performance Leap: Intelligence Index Ties Top Models

According to Artificial Analysis, Alibaba's newly released Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point improvement over its predecessor Qwen3.7 Max's 46, tying Claude Opus 4.8 (max, 56), and trailing only Kimi K3 (max, 57) among Chinese labs, ahead of GLM-5.2 (max, 51).

On GDPval-AA, Qwen3.8 Max achieves 1739 Elo, a 468-point improvement over the previous generation, surpassing Kimi K3 (1685) and roughly matching Claude Fable 5 (1743) and GPT-5.6 Sol (max, 1730), trailing only Claude Opus 5 (max, 1852).

Compared to its predecessor, Qwen3.8 Max shows improvements in agentic evaluations, scientific reasoning, and coding: Terminal-Bench v2.1 up 6 points, CritPt up 7, SciCode up 4, HLE up 3, while GPQA remains unchanged.

Cost and Efficiency: Cost per Task Doubles, but Token Prices Drop

Qwen3.8 Max costs $1.14 per Intelligence Index task, more than double the previous Qwen3.7 Max ($0.53), about 1.3 times Kimi K3 ($0.86), and 2 times GLM-5.2 ($0.57).

The cost increase is mainly due to more rounds in agentic evaluations: GDPval-AA input tokens increase about 15-fold, and output token usage rises 45% to 145M.

Despite higher cost per task, token prices have actually decreased: input/output prices drop from $2.50/$7.50 to $2.00/$6.00, and cache hit price from $0.50 to $0.25.

Concerns: Rising Hallucination Rate and Evaluation Anomalies

On AA-Omniscience, Qwen3.8 Max's score drops 10 points (from +14 to +4), reversing the previous generation's abstention gains. Accuracy remains roughly flat at about 31%, but hallucination rate rises from 23% to 40%, indicating the model tends to attempt rather than abstain when unable to answer.

Additionally, the τ³-Bench Banking result (42%) is considered out-of-distribution, a 32-point improvement over the previous generation, making Qwen3.8 Max outperform higher-scoring models on other evaluations; this anomaly warrants attention.

AA-LCR score drops 2 points, also showing regression in some aspects.

Open-Weight Strategy Shift: From Closed to Open Source

Alibaba announced plans to release Qwen3.8 Max weights next week, marking a strategic shift, as its Max models have typically remained proprietary.

Once released, Qwen3.8 Max will become Alibaba's largest open-weight model, about 6 times the size of its previous largest open-source model Qwen3.5 397B, and the second largest open-weight model after Kimi K3 (2.8T).

The model has 2.4T total parameters, about 95B active parameters, a 1M context window, and supports text, image, and video input, with text output.

Evaluation Corrections and Data Reliability

Artificial Analysis notes that their previously published Qwen3.8 Max score of 53 was affected by intermittent endpoint issues; they have re-run all evaluations on Alibaba's public API endpoint, resulting in a final score of 56.

This correction reminds us that third-party evaluation results may be influenced by test environment, and final data should be based on official or re-tested results.

Credibility boundary

This report is primarily based on evaluation data published by Artificial Analysis on X, a third-party analysis organization. Its data and methods have some credibility, but have not been officially confirmed by Alibaba. Some data (such as cost and parameters) come from official Alibaba statements, but weight release plans are still planned statements and subject to uncertainty.

Insight takeaway

Qwen3.8 Max shows significant performance improvements, but rising costs and increased hallucination rates are major drawbacks. The shift in open-weight strategy may reshape the open-source model landscape, but attention should be paid to its actual release and subsequent performance.

Primary report

Artificial Analysis (X)

Primary source