On the GDPval-AA benchmark, Qwen3.8 Max scores 1739, surpassing Kimi K3's 1685 and trailing only Claude Opus 5's 1852.
However, this high score comes at the cost of a longer reasoning chain: Qwen3.8 Max requires 64 steps per task, while Kimi K3 only needs 14; input tokens increase 15-fold because the test resends the full conversation history at each step.
This leads to a significant rise in inference cost: Qwen3.8 Max's cost per task is $1.14, more than double Qwen3.7 Max's $0.53, despite lower token prices (input drops from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25).