Back to feed
News Story
APriority76
THE DECODER
1 sources

Qwen3.8 Max Matches Claude Opus 4.8, but Kimi K3 Still Scores Higher for 25% Less

Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max. The model catches up to Claude Opus 4.8, but Kimi K3 still achieves a higher score at a 25% lower cost.

SynthePulse Insight · AI deep reading

Qwen3.8 Max Matches Claude Opus 4.8, but Kimi K3 Leads with Lower Cost

Version 1 · 1 source

Alibaba's latest Qwen3.8 Max matches Claude Opus 4.8 on the Artificial Analysis Intelligence Index, but Kimi K3 still leads with a 25% cost advantage. However, Qwen3.8 Max's inference cost has doubled, and it shows regressions such as increased hallucination rates.

  • Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, up 10 points from its predecessor Qwen3.7 Max's 46, tying with Claude Opus 4.8 but trailing Kimi K3's 57.
  • Kimi K3's cost per task is $0.86, about 25% lower than Qwen3.8 Max's $1.14, and it scores higher.
  • On the GDPval-AA benchmark, Qwen3.8 Max scores 1739, surpassing Kimi K3's 1685, but requires 64 steps instead of 14, with input tokens increasing 15-fold.
  • Qwen3.8 Max's AA-LCR drops by 2 points, AA-Omniscience by 10 points, and hallucination rate rises from 23% to 40%.
Open section navigationIntelligence Index: Tied but Not Surpassed

Intelligence Index: Tied but Not Surpassed

According to Artificial Analysis data, Qwen3.8 Max scores 56 on the Intelligence Index, up 10 points from its predecessor Qwen3.7 Max's 46, tying with Claude Opus 4.8 but trailing Kimi K3's 57.

Kimi K3's cost per task is $0.86, about 25% lower than Qwen3.8 Max's $1.14, and it scores higher, showing stronger cost-effectiveness.

GDPval-AA: The Cost Behind the High Score

On the GDPval-AA benchmark, Qwen3.8 Max scores 1739, surpassing Kimi K3's 1685 and trailing only Claude Opus 5's 1852.

However, this high score comes at the cost of a longer reasoning chain: Qwen3.8 Max requires 64 steps per task, while Kimi K3 only needs 14; input tokens increase 15-fold because the test resends the full conversation history at each step.

This leads to a significant rise in inference cost: Qwen3.8 Max's cost per task is $1.14, more than double Qwen3.7 Max's $0.53, despite lower token prices (input drops from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25).

Regressions and Rising Hallucination Rates

Compared to its predecessor, Qwen3.8 Max drops 2 points on AA-LCR, which tests the model's ability to extract information from long texts; AA-Omniscience drops 10 points, which measures the model's accuracy in answering knowledge questions or honestly admitting ignorance.

Accuracy remains around 31%, but hallucination rate jumps from 23% to 40%, indicating the model guesses more often rather than admitting it doesn't know.

Credibility boundary

The information in this article is primarily sourced from Artificial Analysis benchmark tests, as reported by THE DECODER. All data are as claimed by the source and have not been independently verified.

Insight takeaway

Qwen3.8 Max matches Claude Opus 4.8 on the Intelligence Index, but Kimi K3 offers a higher score at a lower cost. However, Qwen3.8 Max's inference cost has doubled, and it shows regressions such as increased hallucination rates, indicating that performance gains come with trade-offs.

Primary report

THE DECODER

Primary source