Back to feed
News Story
THE DECODER
1 sources

Moonshot's Kimi K3 beats Fable 5 in frontend code but lags in complex math

Moonshot's Kimi K3 tops the Code Arena: Frontend rankings, outperforming Claude Fable 5 and GPT-5.6 Sol in frontend code generation. However, it scores only about 39% on FrontierMath Tier 4, far behind OpenAI and Anthropic models that achieve nearly 90%, highlighting a significant gap in advanced math capabilities.

SynthePulse Insight · AI deep reading

Kimi K3: Front-End Code Tops, Complex Math Still Lags — The Real Gap in Chinese AI Models

Version 1 · 1 source

Moonshot's Kimi K3 surpasses Claude Fable 5 and GPT-5.6 Sol on the Code Arena front-end coding benchmark, becoming the first Chinese model to top the leaderboard; but on Epoch AI's FrontierMath Tier 4 expert-level math tasks, it achieves only about 39% accuracy, far below the nearly 90% of top Western models. This contrast reveals both breakthroughs and systemic shortcomings in current Chinese AI models.

  • Kimi K3 ranks first on the Code Arena: Frontend benchmark with a score of 1679, surpassing Claude Fable 5 (1631) and GPT-5.6 Sol (1618), marking the first time a Chinese model has topped this leaderboard.
  • On Epoch AI's FrontierMath Tier 4 (the hardest expert-level math tasks), Kimi K3 achieves only about 39% accuracy, while OpenAI and Anthropic models approach 90%.
  • Kimi K3 excels in code generation but shows a significant gap in complex mathematical reasoning, indicating an uneven capability distribution.
Open section navigationFront-End Code: A First for Chinese Models

Front-End Code: A First for Chinese Models

Moonshot's Kimi K3 scored 1679 on the Code Arena: Frontend benchmark, defeating Claude Fable 5 (1631) and GPT-5.6 Sol (1618), becoming the first Chinese model to rank first on this leaderboard. The benchmark is based on human preference ratings, and Kimi K3 leads all other tested models by a wide margin.

This result shows that in specific application areas such as front-end code generation, Chinese AI models can already match or even surpass top Western models. However, whether this advantage generalizes to other coding tasks remains uncertain.

Complex Math: The Gap Remains Large

According to Epoch AI data, on FrontierMath Tier 4 (the hardest expert-level math tasks in that benchmark), Kimi K3 achieves only about 39% accuracy. In contrast, OpenAI and Anthropic models approach 90% in some cases.

This large gap highlights Kimi K3's clear weakness in complex mathematical reasoning. Despite its strong performance on front-end code, there remains a systemic gap between Chinese models and top Western models on math tasks requiring deep reasoning.

Mixed Performance Reveals Uneven Capability Distribution

Kimi K3's mixed performance reveals an uneven distribution of capabilities in current AI models. In front-end code, Kimi K3 tops human preference ratings, indicating an advantage in generating code that aligns with human preferences; but on math tasks requiring rigorous logical reasoning, its performance lags far behind Western models.

This capability distribution may stem from differences in training data focus or model architecture. Kimi K3's success in code generation may result from targeted optimization, while its deficiency in math reasoning suggests room for improvement in general reasoning ability.

Credibility boundary

This article is based on a report from THE DECODER, which cites public data from Code Arena: Frontend and Epoch AI. All data points come from third-party benchmarks, but no original data links or detailed methodologies are provided. Kimi K3's math score is approximately 39%, and Western models' scores are approximately 90%, both approximate values.

Insight takeaway

Kimi K3 surpasses top Western models in front-end code but lags significantly in complex math, indicating that Chinese AI models have achieved breakthroughs in specific areas but still need to catch up in general reasoning ability.

Primary report

THE DECODER

Primary source